Institutional Learning Agent Architectures
Document: INSTITUTIONAL_LEARNING_AGENT_ARCHITECTURES.md Version: 1.0.0 Date: January 26, 2026 Author: Claude Research Agent Branch: claude/research-agent-architectures-e4Ttd
Executive Summary
This document synthesizes emerging patterns where autonomous research agents evolve from episodic automation into persistent organizational intelligence—systems that accumulate judgment, normalize learning, and compound advantage daily.
These patterns represent a distinct class of architectures not previously covered in MIZOKI's implementation: how autonomous research agents become institutions rather than tools.
1. Core Paradigm Shift: From Automation to Institution
Traditional Automation Model
Pain → Research → Mitigation → Done
↓
(forgotten)
Institutional Learning Model
Pain → Research → Mitigation → Outcome
↓ ↓ ↓ ↓
Curriculum Skill Judgment Learning
↓ ↓ ↓ ↓
└─────────────────────────────────┘
↓
Compounding Organizational Intelligence
Key Difference: Every operation contributes to long-term capability development, not just immediate problem resolution.
2. Pattern 1: Pain Points as Curriculum
Concept
Daily pain points are treated as a learning curriculum rather than a backlog to clear. Systems classify incoming issues by:
| Dimension | Purpose | Learning Signal |
|---|---|---|
| Novelty | New class vs known pattern | Expands problem space coverage |
| Skill Exercised | Diagnosis, causality, optimization, coordination | Builds specific competencies |
| Institutional Weakness | What gap allowed this pain to occur? | Identifies systemic improvements |
| Difficulty Level | Simple fix vs complex investigation | Enables curriculum sequencing |
Current MIZOKI Implementation
Status: ✅ Foundation Complete, ⚠️ Curriculum Structure Missing
# research_ops_contract.py - Current Implementation
class PainPointNormalizer:
"""Ingests 8+ pain sources with priority scoring"""
def compute_priority(self, pain: NormalizedPainPoint) -> float:
return (
self.normalize(pain.urgency) *
pain.impact *
pain.confidence /
(1 + pain.effort)
)
Gap: Priority scoring optimizes for immediate ROI, not learning yield.
Proposed Enhancement: Curriculum-Aware Prioritization
@dataclass
class CurriculumPainPoint(NormalizedPainPoint):
"""Pain point with curriculum metadata"""
# Curriculum dimensions
novelty_score: float = 0.0 # 0=seen before, 1=entirely new
skill_category: SkillType = None # What capability does this exercise?
difficulty_level: int = 1 # 1-5 difficulty scale
prerequisite_skills: List[str] = field(default_factory=list)
# Institutional learning signals
institutional_gap: Optional[str] = None # What systemic weakness exposed?
transfer_potential: float = 0.0 # How applicable to other domains?
def curriculum_priority(self, agent_competence: Dict[str, float]) -> float:
"""Prioritize for optimal learning, not just resolution"""
# Base priority for business value
base = self.compute_priority()
# Learning yield bonus
skill_gap = 1.0 - agent_competence.get(self.skill_category, 0.0)
learning_bonus = self.novelty_score * 0.3 + skill_gap * 0.3
# Curriculum sequencing: Don't attempt too-hard problems
readiness = self._check_prerequisites(agent_competence)
# Transfer bonus: Problems that teach generalizable lessons
transfer_bonus = self.transfer_potential * 0.2
return base * readiness + learning_bonus + transfer_bonus
Curriculum Sequencing Algorithm
class CurriculumSequencer:
"""Sequences pain points for optimal learning progression"""
def sequence_for_learning(
self,
pain_points: List[CurriculumPainPoint],
agent_id: str,
max_batch: int = 5
) -> List[CurriculumPainPoint]:
"""
Select pain points that:
1. Are within agent's current competence (readiness)
2. Stretch capabilities without overwhelming (zone of proximal development)
3. Build toward harder problems (scaffolding)
4. Cover diverse skill categories (breadth)
"""
competence = self.get_agent_competence(agent_id)
# Partition by readiness
ready = [p for p in pain_points if p.difficulty_level <= competence.max_difficulty + 1]
# Score by curriculum value
scored = [(p, p.curriculum_priority(competence.skills)) for p in ready]
scored.sort(key=lambda x: x[1], reverse=True)
# Ensure skill diversity in batch
selected = []
skills_covered = set()
for pain, score in scored:
if len(selected) >= max_batch:
break
if pain.skill_category not in skills_covered or len(skills_covered) >= 3:
selected.append(pain)
skills_covered.add(pain.skill_category)
return selected
3. Pattern 2: Role Rotation and Skill Accretion
Concept
Agents rotate roles across research loops rather than maintaining fixed identities:
| Traditional | Institutional |
|---|---|
| Agent A always synthesizes | Agent A synthesizes today, critiques tomorrow |
| Fixed specialization | Dynamic capability building |
| Local expertise, global blind spots | Cross-pollinated judgment |
Current MIZOKI Implementation
Status: ✅ Role Diversity Implemented, ⚠️ No Rotation Logic
# mizoki_learning_stability.py - Current Implementation
class AgentMarketOrchestrator:
"""10 distinct roles with market-based allocation"""
roles = [
CAUSAL_MODELER, SYSTEMS_DESIGNER, COST_ANALYST,
DATA_SCIENTIST, CODE_ARCHITECT, CREATIVE_STRATEGIST,
RISK_ASSESSOR, DOMAIN_EXPERT, SYNTHESIZER, CRITIC
]
Gap: Agents are assigned to roles based on current competence, not rotated for skill development.
Proposed Enhancement: Rotation-Based Skill Accretion
@dataclass
class AgentRoleHistory:
"""Tracks agent's role assignments over time"""
agent_id: str
role_assignments: List[RoleAssignment] = field(default_factory=list)
skill_levels: Dict[AgentRole, float] = field(default_factory=dict)
rotation_debt: Dict[AgentRole, int] = field(default_factory=dict) # Days since last assignment
class RoleRotationPolicy:
"""Enforces skill accretion through deliberate rotation"""
min_rotation_interval_days: int = 7 # Min days before same role
max_role_concentration: float = 0.4 # Max % of time in single role
skill_decay_rate: float = 0.02 # Daily decay for unused skills
def select_role(
self,
agent: AgentRoleHistory,
task: ResearchTask,
available_roles: List[AgentRole]
) -> AgentRole:
"""
Balance:
1. Task fit (competence for this specific task)
2. Skill development (roles agent hasn't practiced)
3. Rotation debt (roles agent hasn't done recently)
"""
scores = {}
for role in available_roles:
# Task fit
fit = self._compute_task_fit(agent, role, task)
# Skill development need (inverse of current skill)
development_need = 1.0 - agent.skill_levels.get(role, 0.0)
# Rotation debt (log scale for urgency)
debt = agent.rotation_debt.get(role, 0)
rotation_urgency = math.log1p(debt / self.min_rotation_interval_days)
# Weighted combination
scores[role] = (
fit * 0.5 + # Still need competence
development_need * 0.3 + # Build new skills
rotation_urgency * 0.2 # Prevent stagnation
)
return max(scores, key=scores.get)
def update_after_task(
self,
agent: AgentRoleHistory,
role: AgentRole,
outcome: TaskOutcome
) -> None:
"""Update skill levels based on outcome"""
current = agent.skill_levels.get(role, 0.5)
# Learning rate inversely proportional to skill level
# (easier to improve when novice, harder when expert)
learning_rate = 0.1 * (1.0 - current)
# Skill delta based on outcome
if outcome.success:
delta = learning_rate * outcome.difficulty / 5.0
else:
# Still learn from failure, but less
delta = learning_rate * 0.3 * outcome.difficulty / 5.0
agent.skill_levels[role] = min(1.0, current + delta)
# Reset rotation debt for this role
agent.rotation_debt[role] = 0
# Increment debt for other roles
for other_role in agent.skill_levels:
if other_role != role:
agent.rotation_debt[other_role] = agent.rotation_debt.get(other_role, 0) + 1
# Apply decay to all skills not exercised
for other_role in agent.skill_levels:
if other_role != role:
agent.skill_levels[other_role] *= (1.0 - self.skill_decay_rate)
4. Pattern 3: Evaluation Loops that Measure Learning Yield
Concept
Systems measure how much was learned, not just what was fixed:
| Traditional Metric | Learning Yield Metric |
|---|---|
| Issue resolved ✓/✗ | New causal links discovered |
| Time to resolution | Uncertainty reduction for future |
| User satisfaction | Insight reusability across domains |
A mitigation that barely improves metrics but reveals a deep systemic insight may score higher than a clean but shallow fix.
Current MIZOKI Implementation
Status: ✅ Multi-tier Evaluation, ⚠️ No Learning Yield Component
# autonomous_research_patterns_v2.py - Current Implementation
class EvaluationTiers:
"""3-tier evaluation: Static (30%), Predictive (40%), Behavioral (30%)"""
def compute_score(self, proposal, analogs, outcomes):
static = self.static_score(proposal) # 30%
predictive = self.predictive_score(analogs) # 40%
behavioral = self.behavioral_score(outcomes) # 30%
return static * 0.3 + predictive * 0.4 + behavioral * 0.3
Gap: Evaluates proposal quality, not knowledge gained from the research process.
Proposed Enhancement: Learning Yield Evaluation
@dataclass
class LearningYieldMetrics:
"""Measures knowledge gained from research, not just outcome"""
# Causal discovery
new_causal_links: int = 0 # Novel cause-effect relationships found
invalidated_hypotheses: int = 0 # Wrong beliefs corrected
uncertainty_reduction: float = 0.0 # Entropy decrease in KG
# Insight quality
generalizability: float = 0.0 # Applicable to other domains
mechanism_depth: int = 0 # Levels of "why" explained
novelty: float = 0.0 # Not predicted by prior knowledge
# Reusability
artifact_reuse_potential: float = 0.0 # Can outputs help future research?
pattern_extracted: Optional[str] = None # Named pattern for future reference
def compute_learning_yield(self) -> float:
"""Composite learning yield score"""
discovery = (
self.new_causal_links * 0.3 +
self.invalidated_hypotheses * 0.2 +
self.uncertainty_reduction * 0.5
)
quality = (
self.generalizability * 0.4 +
self.mechanism_depth / 5.0 * 0.3 +
self.novelty * 0.3
)
reuse = (
self.artifact_reuse_potential * 0.6 +
(1.0 if self.pattern_extracted else 0.0) * 0.4
)
return discovery * 0.4 + quality * 0.35 + reuse * 0.25
class LearningYieldEvaluator:
"""Evaluates research for learning yield, not just outcome"""
def evaluate_research_outcome(
self,
research_session: ResearchSession,
outcome: MitigationOutcome,
kg_before: KGSnapshot,
kg_after: KGSnapshot
) -> LearningYieldMetrics:
"""Compute learning yield from research"""
metrics = LearningYieldMetrics()
# Causal discovery: Count new edges
new_edges = kg_after.edges - kg_before.edges
metrics.new_causal_links = len([
e for e in new_edges
if e.edge_type in ['CAUSED_BY', 'MITIGATED_BY', 'INVALIDATED_BY']
])
# Invalidated hypotheses: Count edges removed or marked invalid
metrics.invalidated_hypotheses = len([
h for h in research_session.hypotheses_tested
if h.status == HypothesisStatus.INVALIDATED
])
# Uncertainty reduction: KG entropy change
metrics.uncertainty_reduction = self._compute_entropy_reduction(
kg_before, kg_after
)
# Generalizability: Does insight apply beyond this specific case?
metrics.generalizability = self._assess_generalizability(
research_session.insights,
domain=research_session.domain
)
# Mechanism depth: How many levels of "why" explained?
metrics.mechanism_depth = self._count_causal_depth(
research_session.causal_chain
)
# Reusability: Rate artifacts for future use
metrics.artifact_reuse_potential = self._rate_artifact_reuse(
research_session.artifacts
)
return metrics
def _compute_entropy_reduction(
self,
before: KGSnapshot,
after: KGSnapshot
) -> float:
"""
Compute reduction in KG uncertainty.
Higher confidence edges = lower entropy.
"""
def kg_entropy(snapshot: KGSnapshot) -> float:
# Entropy based on edge confidence distribution
confidences = [e.confidence for e in snapshot.edges]
if not confidences:
return 1.0 # Max uncertainty
# Shannon entropy of confidence distribution
return -sum(
p * math.log2(p) if p > 0 else 0
for p in self._normalize(confidences)
)
before_entropy = kg_entropy(before)
after_entropy = kg_entropy(after)
# Reduction (positive = good)
return max(0, before_entropy - after_entropy)
Composite Evaluation with Learning Yield
class InstitutionalEvaluator:
"""Combines task performance with learning yield"""
def evaluate_complete(
self,
outcome: MitigationOutcome,
learning_yield: LearningYieldMetrics
) -> InstitutionalScore:
"""
Score = Task Performance × Learning Multiplier
A mediocre fix that teaches a lot can score higher than
a perfect fix that teaches nothing.
"""
# Task performance (existing 3-tier evaluation)
task_score = self.task_evaluator.compute_score(outcome)
# Learning yield
yield_score = learning_yield.compute_learning_yield()
# Learning multiplier (1.0 to 1.5)
learning_multiplier = 1.0 + yield_score * 0.5
# Composite
composite = task_score * learning_multiplier
return InstitutionalScore(
task_score=task_score,
learning_yield=yield_score,
composite=composite,
recommendation=self._generate_recommendation(task_score, yield_score)
)
def _generate_recommendation(
self,
task_score: float,
learning_yield: float
) -> str:
"""Generate actionable recommendation"""
if task_score >= 0.8 and learning_yield >= 0.7:
return "EXEMPLAR: Document pattern for reuse"
elif task_score >= 0.8 and learning_yield < 0.3:
return "ROUTINE: Efficient but low insight, consider automation"
elif task_score < 0.5 and learning_yield >= 0.7:
return "VALUABLE_FAILURE: Rich learning despite poor outcome"
elif task_score < 0.5 and learning_yield < 0.3:
return "POOR: Review process, consider escalation"
else:
return "NORMAL: Standard outcome with moderate learning"
5. Pattern 4: Knowledge Graph as Institutional Brain
Concept
Knowledge graphs become the primary locus of intelligence, encoding not just facts but:
| Traditional KG | Institutional KG |
|---|---|
| Entity relationships | Decision rationales |
| Current state | Trade-offs considered |
| Facts | Why options were NOT chosen |
| What we know | How we reason |
Query examples: - "How would our past selves think about this problem?" - "What mistakes did we almost make in similar situations?" - "What assumptions have we invalidated recently?"
Current MIZOKI Implementation
Status: ✅ KG Brain Implemented, ⚠️ Reasoning History Not Captured
# knowledge_graph_brain_integration.py - Current Implementation
class KGNode:
"""Basic node with embeddings and properties"""
node_id: str
node_type: str
embeddings: Optional[List[float]]
properties: Dict[str, Any]
class KGEdge:
"""Edge with confidence and provenance"""
edge_id: str
source_id: str
target_id: str
edge_type: str
confidence: float
provenance: Optional[str]
Gap: Captures facts and relationships, but not reasoning history, rejected alternatives, or decision context.
Proposed Enhancement: Institutional Memory KG
@dataclass
class ReasoningHistoryNode(KGNode):
"""Captures HOW the organization reasons, not just WHAT it knows"""
# Decision context
decision_id: str # Link to original decision
decision_timestamp: datetime
decision_context: Dict[str, Any] # Full context at decision time
# Options considered
options_evaluated: List[DecisionOption]
selected_option: str
selection_rationale: str
# Rejected alternatives
rejected_options: List[RejectedOption]
rejection_reasons: Dict[str, str]
# Confidence evolution
initial_confidence: float
final_confidence: float
confidence_factors: List[str]
# Meta-information
reasoning_paradigm: str # CoT, ToT, GoT, etc.
key_assumptions: List[str]
assumption_validations: Dict[str, bool]
@dataclass
class RejectedOption:
"""Records why an option was NOT chosen"""
option_id: str
option_description: str
rejection_reason: str
rejection_confidence: float
conditions_for_reconsideration: Optional[str] # When might this become viable?
@dataclass
class InstitutionalMemoryEdge(KGEdge):
"""Edge that captures reasoning relationship"""
# Reasoning metadata
discovered_how: str # How was this relationship discovered?
discovery_confidence: float # How confident at discovery?
current_confidence: float # Current confidence after validation
validation_history: List[ValidationEvent]
# Counterfactual information
if_not_true: str # What would be different if this edge were false?
invalidation_conditions: List[str] # Conditions that would invalidate this
# Usage tracking
times_used_in_reasoning: int
times_validated_by_outcome: int
times_contradicted_by_outcome: int
class InstitutionalMemoryKG:
"""KG that captures organizational reasoning, not just facts"""
def record_decision(
self,
decision: Decision,
context: DecisionContext,
options: List[DecisionOption],
selected: str,
rejected: List[RejectedOption]
) -> ReasoningHistoryNode:
"""Record a decision with full reasoning context"""
node = ReasoningHistoryNode(
node_id=f"decision:{decision.id}",
node_type="reasoning_history",
decision_id=decision.id,
decision_timestamp=datetime.utcnow(),
decision_context=context.to_dict(),
options_evaluated=options,
selected_option=selected,
selection_rationale=decision.rationale,
rejected_options=rejected,
rejection_reasons={r.option_id: r.rejection_reason for r in rejected},
initial_confidence=decision.initial_confidence,
final_confidence=decision.final_confidence,
reasoning_paradigm=context.paradigm,
key_assumptions=decision.assumptions
)
self.add_node(node)
# Link to relevant entity nodes
for entity_id in decision.affected_entities:
self.add_edge(InstitutionalMemoryEdge(
source_id=node.node_id,
target_id=entity_id,
edge_type="DECIDED_FOR",
discovered_how="decision_recording"
))
return node
def query_past_reasoning(
self,
problem_description: str,
similarity_threshold: float = 0.7
) -> List[ReasoningHistoryNode]:
"""
Query: "How would our past selves think about this problem?"
Returns similar past decisions with full reasoning context.
"""
# Embed the problem
problem_embedding = self.embedding_service.embed(problem_description)
# Find similar reasoning history nodes
similar_nodes = []
for node in self.get_nodes_by_type("reasoning_history"):
context_embedding = self.embedding_service.embed(
str(node.decision_context)
)
similarity = cosine_similarity(problem_embedding, context_embedding)
if similarity >= similarity_threshold:
similar_nodes.append((node, similarity))
# Sort by similarity
similar_nodes.sort(key=lambda x: x[1], reverse=True)
return [node for node, _ in similar_nodes[:10]]
def query_rejected_options(
self,
current_context: DecisionContext
) -> List[RejectedOption]:
"""
Query: "What options have we rejected in similar situations?"
Returns rejected options that might be worth reconsidering.
"""
past_decisions = self.query_past_reasoning(str(current_context))
reconsidered = []
for decision in past_decisions:
for rejected in decision.rejected_options:
# Check if conditions for reconsideration are now met
if rejected.conditions_for_reconsideration:
if self._check_conditions_met(
rejected.conditions_for_reconsideration,
current_context
):
reconsidered.append(rejected)
return reconsidered
def query_invalidated_assumptions(
self,
domain: str,
lookback_days: int = 30
) -> List[Tuple[str, str]]:
"""
Query: "What assumptions have we invalidated recently?"
Returns (assumption, invalidation_reason) pairs.
"""
cutoff = datetime.utcnow() - timedelta(days=lookback_days)
invalidated = []
for node in self.get_nodes_by_type("reasoning_history"):
if node.decision_timestamp < cutoff:
continue
for assumption, validated in node.assumption_validations.items():
if not validated:
invalidated.append((
assumption,
f"Invalidated in decision {node.decision_id}"
))
return invalidated
6. Pattern 5: Continuous Improvement via Meta-Learning Loops
Concept
Systems improve not just outcomes, but:
| Level | What Improves |
|---|---|
| Object-level | Individual decisions |
| Process-level | How decisions are made |
| Meta-level | How the process improves |
Meta-learning loops periodically audit: - Research loop structure - Evidence standards - Coordination patterns - Evaluation thresholds
Current MIZOKI Implementation
Status: ✅ Dual-horizon Cycles, ⚠️ No Process Improvement Loop
# mizoki_learning_stability.py - Current Implementation
class DualHorizonCycles:
"""Tactical (24h) + Strategic (7d) cycles"""
tactical_metrics = ['pain_points_addressed', 'mitigations_proposed']
strategic_metrics = ['kg_refactoring', 'hypothesis_invalidation']
Gap: Cycles improve outcomes, but don't improve the research process itself.
Proposed Enhancement: Meta-Learning Controller
@dataclass
class ProcessMetrics:
"""Metrics about the research process, not outcomes"""
# Research efficiency
avg_time_to_hypothesis: timedelta
avg_hypotheses_per_pain: float
hypothesis_validation_rate: float
# Coordination quality
role_handoff_smoothness: float # 0-1, higher = cleaner handoffs
agent_utilization: float # % of agent capacity used
coordination_overhead: float # Time spent coordinating vs researching
# Evidence quality
avg_evidence_depth: float # Layers of supporting evidence
evidence_contradiction_rate: float # % of evidence later contradicted
evidence_reuse_rate: float # % of evidence reused in other research
# Evaluation calibration
prediction_accuracy: float # How well predictions match outcomes
confidence_calibration: float # Is stated confidence accurate?
escalation_appropriateness: float # Were escalations necessary/sufficient?
class MetaLearningController:
"""Improves the research process itself, not just outcomes"""
audit_interval_days: int = 7
adjustment_rate: float = 0.05 # 5% max change per audit
min_samples_for_adjustment: int = 20
def audit_process(
self,
period_start: datetime,
period_end: datetime
) -> ProcessAuditResult:
"""Audit research process for improvement opportunities"""
# Collect process metrics
metrics = self._compute_process_metrics(period_start, period_end)
# Compare to baselines
baselines = self._get_baselines()
deviations = self._compute_deviations(metrics, baselines)
# Identify improvement opportunities
opportunities = []
# Research efficiency
if metrics.avg_hypotheses_per_pain > 5:
opportunities.append(ProcessImprovement(
area="hypothesis_generation",
issue="Too many hypotheses generated",
recommendation="Tighten initial filtering criteria",
adjustment={"hypothesis_confidence_threshold": +0.05}
))
# Coordination quality
if metrics.coordination_overhead > 0.3:
opportunities.append(ProcessImprovement(
area="coordination",
issue="High coordination overhead",
recommendation="Reduce handoff frequency, increase agent autonomy",
adjustment={"autonomy_threshold": -0.05}
))
# Evidence quality
if metrics.evidence_contradiction_rate > 0.2:
opportunities.append(ProcessImprovement(
area="evidence",
issue="High evidence contradiction rate",
recommendation="Increase evidence validation before use",
adjustment={"evidence_validation_depth": +1}
))
# Evaluation calibration
if metrics.confidence_calibration < 0.7:
opportunities.append(ProcessImprovement(
area="evaluation",
issue="Poor confidence calibration",
recommendation="Adjust confidence scoring formula",
adjustment={"confidence_dampening_factor": +0.1}
))
return ProcessAuditResult(
period=(period_start, period_end),
metrics=metrics,
deviations=deviations,
improvements=opportunities
)
def apply_process_improvements(
self,
audit_result: ProcessAuditResult
) -> List[ProcessAdjustment]:
"""Apply approved process improvements"""
adjustments = []
for improvement in audit_result.improvements:
# Check if we have enough samples
samples = self._count_relevant_samples(improvement.area)
if samples < self.min_samples_for_adjustment:
continue
# Apply adjustment with rate limiting
for param, delta in improvement.adjustment.items():
current = self.get_param(param)
new_value = current + delta * self.adjustment_rate
# Validate bounds
new_value = self._clamp_to_bounds(param, new_value)
self.set_param(param, new_value)
adjustments.append(ProcessAdjustment(
param=param,
old_value=current,
new_value=new_value,
reason=improvement.recommendation
))
return adjustments
def audit_meta_process(self) -> MetaAuditResult:
"""Audit the auditing process itself (meta-meta-learning)"""
# How effective have our process improvements been?
recent_audits = self._get_recent_audits(n=10)
improvement_effectiveness = []
for audit in recent_audits:
for adjustment in audit.applied_adjustments:
# Compare metrics before and after adjustment
before = self._get_metrics_at(audit.period[0])
after = self._get_metrics_at(audit.period[1])
effectiveness = self._compute_adjustment_effectiveness(
adjustment, before, after
)
improvement_effectiveness.append(effectiveness)
# Meta-recommendation
avg_effectiveness = sum(improvement_effectiveness) / len(improvement_effectiveness)
if avg_effectiveness < 0.3:
meta_recommendation = "Process improvements not effective; consider fundamental restructure"
elif avg_effectiveness < 0.6:
meta_recommendation = "Process improvements partially effective; tune adjustment rate"
else:
meta_recommendation = "Process improvements effective; maintain current approach"
return MetaAuditResult(
improvement_effectiveness=avg_effectiveness,
recommendation=meta_recommendation,
suggested_adjustment_rate=self._suggest_adjustment_rate(avg_effectiveness)
)
7. Implementation Roadmap for MIZOKI
Phase 1: Close Learning Feedback Loop (Week 1-2)
Objective: Wire learning signals back into decision-making
# Integration point in research_ops_contract.py
class EnhancedResearchOrchestrator:
def __init__(self):
self.learning_feedback = LearningFeedbackLoop()
async def process_pain_point(self, pain: NormalizedPainPoint):
# Enrich with curriculum metadata
curriculum_pain = self.learning_feedback.enrich_with_curriculum(pain)
# Prioritize with learning yield
priority = curriculum_pain.curriculum_priority(
self.learning_feedback.get_agent_competence()
)
# Continue with existing flow...
Deliverables:
- [ ] LearningFeedbackLoop class
- [ ] Integration with PainPointNormalizer
- [ ] Integration with AgentMarketOrchestrator
- [ ] Firestore schema for competence tracking
Phase 2: Institutional Memory KG (Week 3-4)
Objective: Capture reasoning history, not just facts
# Integration in knowledge_graph_brain_integration.py
class EnhancedKGBrain:
def record_decision_with_history(self, decision: Decision):
# Existing: Record decision
self.record_decision(decision)
# New: Record reasoning history
self.institutional_memory.record_decision(
decision=decision,
context=self.current_context,
options=decision.options_evaluated,
selected=decision.selected,
rejected=decision.rejected_options
)
Deliverables:
- [ ] ReasoningHistoryNode schema
- [ ] InstitutionalMemoryKG class
- [ ] Query methods for past reasoning
- [ ] Migration script for existing decisions
Phase 3: Learning Yield Evaluation (Week 5-6)
Objective: Measure knowledge gained, not just outcomes
# Integration in autonomous_research_patterns_v2.py
class EnhancedResearchPipeline:
async def evaluate_complete(self, outcome: MitigationOutcome):
# Existing: Task evaluation
task_score = await self.evaluate_task(outcome)
# New: Learning yield evaluation
learning_yield = await self.learning_evaluator.evaluate_research_outcome(
research_session=self.current_session,
outcome=outcome,
kg_before=self.kg_snapshot_before,
kg_after=self.kg_snapshot_after
)
# Composite score
return self.institutional_evaluator.evaluate_complete(
outcome=outcome,
learning_yield=learning_yield
)
Deliverables:
- [ ] LearningYieldMetrics class
- [ ] LearningYieldEvaluator class
- [ ] InstitutionalEvaluator class
- [ ] Dashboard for learning yield visualization
Phase 4: Meta-Learning Controller (Week 7-8)
Objective: Improve the research process itself
# New module: meta_learning_controller.py
class MetaLearningOrchestrator:
def __init__(self):
self.controller = MetaLearningController()
async def run_weekly_audit(self):
# Audit process metrics
audit = await self.controller.audit_process(
period_start=datetime.utcnow() - timedelta(days=7),
period_end=datetime.utcnow()
)
# Apply improvements
adjustments = await self.controller.apply_process_improvements(audit)
# Log for visibility
await self.log_audit_results(audit, adjustments)
return audit
Deliverables:
- [ ] MetaLearningController class
- [ ] Process metrics collection
- [ ] Weekly audit job
- [ ] Process improvement application
Phase 5: Role Rotation Policy (Week 9-10)
Objective: Build skills through deliberate rotation
# Integration in mizoki_learning_stability.py
class EnhancedAgentMarketOrchestrator:
def __init__(self):
self.rotation_policy = RoleRotationPolicy()
def assign_role(self, agent_id: str, task: ResearchTask):
# Get agent history
history = self.get_agent_history(agent_id)
# Select role with rotation consideration
role = self.rotation_policy.select_role(
agent=history,
task=task,
available_roles=self.available_roles
)
return role
Deliverables:
- [ ] RoleRotationPolicy class
- [ ] AgentRoleHistory tracking
- [ ] Skill decay mechanism
- [ ] Rotation debt tracking
8. Success Metrics
Institutional Learning KPIs
| Metric | Baseline | Target | Measurement |
|---|---|---|---|
| Learning Yield per Research | N/A | 0.5+ | Average LearningYieldMetrics.compute_learning_yield() |
| Knowledge Reuse Rate | N/A | 30%+ | % of insights reused in subsequent research |
| Agent Skill Breadth | 3 roles | 6+ roles | Avg roles per agent with skill > 0.5 |
| Reasoning History Coverage | 0% | 80%+ | % of decisions with full reasoning context |
| Process Improvement Effectiveness | N/A | 60%+ | % of process changes that improve metrics |
| Curriculum Completion | N/A | 70%+ | % of skill categories exercised per quarter |
Compound Advantage Indicators
| Indicator | Description | Target Trend |
|---|---|---|
| Time to Resolution | Avg time to resolve similar pain types | ↓ 20% quarterly |
| First-attempt Success | % of mitigations that work on first try | ↑ 10% quarterly |
| Novel Pain Handling | Success rate on never-before-seen problems | ↑ 5% quarterly |
| Cross-Domain Transfer | Success when applying lessons from other domains | ↑ 15% quarterly |
9. Conclusion
The patterns documented here represent a paradigm shift from autonomous agents as tools to autonomous agents as institutions. By implementing:
- Pain as Curriculum - Every operation contributes to capability building
- Role Rotation - Agents develop cross-functional judgment
- Learning Yield Evaluation - Success measured by knowledge gained
- Institutional KG - Reasoning history persisted for future reference
- Meta-Learning - The process improves itself continuously
MIZOKI can evolve from a sophisticated automation platform into a living organizational mind that remembers, reasons, adapts, and improves continuously—even as people, systems, and priorities change.
Strategic Outcome: Compound organizational intelligence that outpaces competitors through accumulated judgment, not just accumulated data.
References
- MIZOKI Codebase Analysis (January 2026)
- Autonomous Research Patterns V2 (
autonomous_research_patterns_v2.py) - Learning Stability Framework (
mizoki_learning_stability.py) - Research Ops Contract (
research_ops_contract.py) - Knowledge Graph Brain Integration (
knowledge_graph_brain_integration.py)