Institutional Learning Agent Architectures

Document: INSTITUTIONAL_LEARNING_AGENT_ARCHITECTURES.md Version: 1.0.0 Date: January 26, 2026 Author: Claude Research Agent Branch: claude/research-agent-architectures-e4Ttd


Executive Summary

This document synthesizes emerging patterns where autonomous research agents evolve from episodic automation into persistent organizational intelligence—systems that accumulate judgment, normalize learning, and compound advantage daily.

These patterns represent a distinct class of architectures not previously covered in MIZOKI's implementation: how autonomous research agents become institutions rather than tools.


1. Core Paradigm Shift: From Automation to Institution

Traditional Automation Model

Pain → Research → Mitigation → Done
       ↓
    (forgotten)

Institutional Learning Model

Pain → Research → Mitigation → Outcome
  ↓         ↓           ↓          ↓
Curriculum  Skill    Judgment   Learning
   ↓         ↓           ↓          ↓
  └─────────────────────────────────┘
              ↓
    Compounding Organizational Intelligence

Key Difference: Every operation contributes to long-term capability development, not just immediate problem resolution.


2. Pattern 1: Pain Points as Curriculum

Concept

Daily pain points are treated as a learning curriculum rather than a backlog to clear. Systems classify incoming issues by:

Dimension Purpose Learning Signal
Novelty New class vs known pattern Expands problem space coverage
Skill Exercised Diagnosis, causality, optimization, coordination Builds specific competencies
Institutional Weakness What gap allowed this pain to occur? Identifies systemic improvements
Difficulty Level Simple fix vs complex investigation Enables curriculum sequencing

Current MIZOKI Implementation

Status: ✅ Foundation Complete, ⚠️ Curriculum Structure Missing

# research_ops_contract.py - Current Implementation
class PainPointNormalizer:
    """Ingests 8+ pain sources with priority scoring"""

    def compute_priority(self, pain: NormalizedPainPoint) -> float:
        return (
            self.normalize(pain.urgency) *
            pain.impact *
            pain.confidence /
            (1 + pain.effort)
        )

Gap: Priority scoring optimizes for immediate ROI, not learning yield.

Proposed Enhancement: Curriculum-Aware Prioritization

@dataclass
class CurriculumPainPoint(NormalizedPainPoint):
    """Pain point with curriculum metadata"""

    # Curriculum dimensions
    novelty_score: float = 0.0          # 0=seen before, 1=entirely new
    skill_category: SkillType = None    # What capability does this exercise?
    difficulty_level: int = 1           # 1-5 difficulty scale
    prerequisite_skills: List[str] = field(default_factory=list)

    # Institutional learning signals
    institutional_gap: Optional[str] = None  # What systemic weakness exposed?
    transfer_potential: float = 0.0     # How applicable to other domains?

    def curriculum_priority(self, agent_competence: Dict[str, float]) -> float:
        """Prioritize for optimal learning, not just resolution"""

        # Base priority for business value
        base = self.compute_priority()

        # Learning yield bonus
        skill_gap = 1.0 - agent_competence.get(self.skill_category, 0.0)
        learning_bonus = self.novelty_score * 0.3 + skill_gap * 0.3

        # Curriculum sequencing: Don't attempt too-hard problems
        readiness = self._check_prerequisites(agent_competence)

        # Transfer bonus: Problems that teach generalizable lessons
        transfer_bonus = self.transfer_potential * 0.2

        return base * readiness + learning_bonus + transfer_bonus

Curriculum Sequencing Algorithm

class CurriculumSequencer:
    """Sequences pain points for optimal learning progression"""

    def sequence_for_learning(
        self,
        pain_points: List[CurriculumPainPoint],
        agent_id: str,
        max_batch: int = 5
    ) -> List[CurriculumPainPoint]:
        """
        Select pain points that:
        1. Are within agent's current competence (readiness)
        2. Stretch capabilities without overwhelming (zone of proximal development)
        3. Build toward harder problems (scaffolding)
        4. Cover diverse skill categories (breadth)
        """

        competence = self.get_agent_competence(agent_id)

        # Partition by readiness
        ready = [p for p in pain_points if p.difficulty_level <= competence.max_difficulty + 1]

        # Score by curriculum value
        scored = [(p, p.curriculum_priority(competence.skills)) for p in ready]
        scored.sort(key=lambda x: x[1], reverse=True)

        # Ensure skill diversity in batch
        selected = []
        skills_covered = set()

        for pain, score in scored:
            if len(selected) >= max_batch:
                break
            if pain.skill_category not in skills_covered or len(skills_covered) >= 3:
                selected.append(pain)
                skills_covered.add(pain.skill_category)

        return selected

3. Pattern 2: Role Rotation and Skill Accretion

Concept

Agents rotate roles across research loops rather than maintaining fixed identities:

Traditional Institutional
Agent A always synthesizes Agent A synthesizes today, critiques tomorrow
Fixed specialization Dynamic capability building
Local expertise, global blind spots Cross-pollinated judgment

Current MIZOKI Implementation

Status: ✅ Role Diversity Implemented, ⚠️ No Rotation Logic

# mizoki_learning_stability.py - Current Implementation
class AgentMarketOrchestrator:
    """10 distinct roles with market-based allocation"""

    roles = [
        CAUSAL_MODELER, SYSTEMS_DESIGNER, COST_ANALYST,
        DATA_SCIENTIST, CODE_ARCHITECT, CREATIVE_STRATEGIST,
        RISK_ASSESSOR, DOMAIN_EXPERT, SYNTHESIZER, CRITIC
    ]

Gap: Agents are assigned to roles based on current competence, not rotated for skill development.

Proposed Enhancement: Rotation-Based Skill Accretion

@dataclass
class AgentRoleHistory:
    """Tracks agent's role assignments over time"""
    agent_id: str
    role_assignments: List[RoleAssignment] = field(default_factory=list)
    skill_levels: Dict[AgentRole, float] = field(default_factory=dict)
    rotation_debt: Dict[AgentRole, int] = field(default_factory=dict)  # Days since last assignment

class RoleRotationPolicy:
    """Enforces skill accretion through deliberate rotation"""

    min_rotation_interval_days: int = 7    # Min days before same role
    max_role_concentration: float = 0.4    # Max % of time in single role
    skill_decay_rate: float = 0.02         # Daily decay for unused skills

    def select_role(
        self,
        agent: AgentRoleHistory,
        task: ResearchTask,
        available_roles: List[AgentRole]
    ) -> AgentRole:
        """
        Balance:
        1. Task fit (competence for this specific task)
        2. Skill development (roles agent hasn't practiced)
        3. Rotation debt (roles agent hasn't done recently)
        """

        scores = {}
        for role in available_roles:
            # Task fit
            fit = self._compute_task_fit(agent, role, task)

            # Skill development need (inverse of current skill)
            development_need = 1.0 - agent.skill_levels.get(role, 0.0)

            # Rotation debt (log scale for urgency)
            debt = agent.rotation_debt.get(role, 0)
            rotation_urgency = math.log1p(debt / self.min_rotation_interval_days)

            # Weighted combination
            scores[role] = (
                fit * 0.5 +                    # Still need competence
                development_need * 0.3 +       # Build new skills
                rotation_urgency * 0.2         # Prevent stagnation
            )

        return max(scores, key=scores.get)

    def update_after_task(
        self,
        agent: AgentRoleHistory,
        role: AgentRole,
        outcome: TaskOutcome
    ) -> None:
        """Update skill levels based on outcome"""

        current = agent.skill_levels.get(role, 0.5)

        # Learning rate inversely proportional to skill level
        # (easier to improve when novice, harder when expert)
        learning_rate = 0.1 * (1.0 - current)

        # Skill delta based on outcome
        if outcome.success:
            delta = learning_rate * outcome.difficulty / 5.0
        else:
            # Still learn from failure, but less
            delta = learning_rate * 0.3 * outcome.difficulty / 5.0

        agent.skill_levels[role] = min(1.0, current + delta)

        # Reset rotation debt for this role
        agent.rotation_debt[role] = 0

        # Increment debt for other roles
        for other_role in agent.skill_levels:
            if other_role != role:
                agent.rotation_debt[other_role] = agent.rotation_debt.get(other_role, 0) + 1

        # Apply decay to all skills not exercised
        for other_role in agent.skill_levels:
            if other_role != role:
                agent.skill_levels[other_role] *= (1.0 - self.skill_decay_rate)

4. Pattern 3: Evaluation Loops that Measure Learning Yield

Concept

Systems measure how much was learned, not just what was fixed:

Traditional Metric Learning Yield Metric
Issue resolved ✓/✗ New causal links discovered
Time to resolution Uncertainty reduction for future
User satisfaction Insight reusability across domains

A mitigation that barely improves metrics but reveals a deep systemic insight may score higher than a clean but shallow fix.

Current MIZOKI Implementation

Status: ✅ Multi-tier Evaluation, ⚠️ No Learning Yield Component

# autonomous_research_patterns_v2.py - Current Implementation
class EvaluationTiers:
    """3-tier evaluation: Static (30%), Predictive (40%), Behavioral (30%)"""

    def compute_score(self, proposal, analogs, outcomes):
        static = self.static_score(proposal)       # 30%
        predictive = self.predictive_score(analogs) # 40%
        behavioral = self.behavioral_score(outcomes) # 30%
        return static * 0.3 + predictive * 0.4 + behavioral * 0.3

Gap: Evaluates proposal quality, not knowledge gained from the research process.

Proposed Enhancement: Learning Yield Evaluation

@dataclass
class LearningYieldMetrics:
    """Measures knowledge gained from research, not just outcome"""

    # Causal discovery
    new_causal_links: int = 0           # Novel cause-effect relationships found
    invalidated_hypotheses: int = 0     # Wrong beliefs corrected
    uncertainty_reduction: float = 0.0  # Entropy decrease in KG

    # Insight quality
    generalizability: float = 0.0       # Applicable to other domains
    mechanism_depth: int = 0            # Levels of "why" explained
    novelty: float = 0.0                # Not predicted by prior knowledge

    # Reusability
    artifact_reuse_potential: float = 0.0  # Can outputs help future research?
    pattern_extracted: Optional[str] = None # Named pattern for future reference

    def compute_learning_yield(self) -> float:
        """Composite learning yield score"""

        discovery = (
            self.new_causal_links * 0.3 +
            self.invalidated_hypotheses * 0.2 +
            self.uncertainty_reduction * 0.5
        )

        quality = (
            self.generalizability * 0.4 +
            self.mechanism_depth / 5.0 * 0.3 +
            self.novelty * 0.3
        )

        reuse = (
            self.artifact_reuse_potential * 0.6 +
            (1.0 if self.pattern_extracted else 0.0) * 0.4
        )

        return discovery * 0.4 + quality * 0.35 + reuse * 0.25


class LearningYieldEvaluator:
    """Evaluates research for learning yield, not just outcome"""

    def evaluate_research_outcome(
        self,
        research_session: ResearchSession,
        outcome: MitigationOutcome,
        kg_before: KGSnapshot,
        kg_after: KGSnapshot
    ) -> LearningYieldMetrics:
        """Compute learning yield from research"""

        metrics = LearningYieldMetrics()

        # Causal discovery: Count new edges
        new_edges = kg_after.edges - kg_before.edges
        metrics.new_causal_links = len([
            e for e in new_edges
            if e.edge_type in ['CAUSED_BY', 'MITIGATED_BY', 'INVALIDATED_BY']
        ])

        # Invalidated hypotheses: Count edges removed or marked invalid
        metrics.invalidated_hypotheses = len([
            h for h in research_session.hypotheses_tested
            if h.status == HypothesisStatus.INVALIDATED
        ])

        # Uncertainty reduction: KG entropy change
        metrics.uncertainty_reduction = self._compute_entropy_reduction(
            kg_before, kg_after
        )

        # Generalizability: Does insight apply beyond this specific case?
        metrics.generalizability = self._assess_generalizability(
            research_session.insights,
            domain=research_session.domain
        )

        # Mechanism depth: How many levels of "why" explained?
        metrics.mechanism_depth = self._count_causal_depth(
            research_session.causal_chain
        )

        # Reusability: Rate artifacts for future use
        metrics.artifact_reuse_potential = self._rate_artifact_reuse(
            research_session.artifacts
        )

        return metrics

    def _compute_entropy_reduction(
        self,
        before: KGSnapshot,
        after: KGSnapshot
    ) -> float:
        """
        Compute reduction in KG uncertainty.
        Higher confidence edges = lower entropy.
        """

        def kg_entropy(snapshot: KGSnapshot) -> float:
            # Entropy based on edge confidence distribution
            confidences = [e.confidence for e in snapshot.edges]
            if not confidences:
                return 1.0  # Max uncertainty

            # Shannon entropy of confidence distribution
            return -sum(
                p * math.log2(p) if p > 0 else 0
                for p in self._normalize(confidences)
            )

        before_entropy = kg_entropy(before)
        after_entropy = kg_entropy(after)

        # Reduction (positive = good)
        return max(0, before_entropy - after_entropy)

Composite Evaluation with Learning Yield

class InstitutionalEvaluator:
    """Combines task performance with learning yield"""

    def evaluate_complete(
        self,
        outcome: MitigationOutcome,
        learning_yield: LearningYieldMetrics
    ) -> InstitutionalScore:
        """
        Score = Task Performance × Learning Multiplier

        A mediocre fix that teaches a lot can score higher than
        a perfect fix that teaches nothing.
        """

        # Task performance (existing 3-tier evaluation)
        task_score = self.task_evaluator.compute_score(outcome)

        # Learning yield
        yield_score = learning_yield.compute_learning_yield()

        # Learning multiplier (1.0 to 1.5)
        learning_multiplier = 1.0 + yield_score * 0.5

        # Composite
        composite = task_score * learning_multiplier

        return InstitutionalScore(
            task_score=task_score,
            learning_yield=yield_score,
            composite=composite,
            recommendation=self._generate_recommendation(task_score, yield_score)
        )

    def _generate_recommendation(
        self,
        task_score: float,
        learning_yield: float
    ) -> str:
        """Generate actionable recommendation"""

        if task_score >= 0.8 and learning_yield >= 0.7:
            return "EXEMPLAR: Document pattern for reuse"
        elif task_score >= 0.8 and learning_yield < 0.3:
            return "ROUTINE: Efficient but low insight, consider automation"
        elif task_score < 0.5 and learning_yield >= 0.7:
            return "VALUABLE_FAILURE: Rich learning despite poor outcome"
        elif task_score < 0.5 and learning_yield < 0.3:
            return "POOR: Review process, consider escalation"
        else:
            return "NORMAL: Standard outcome with moderate learning"

5. Pattern 4: Knowledge Graph as Institutional Brain

Concept

Knowledge graphs become the primary locus of intelligence, encoding not just facts but:

Traditional KG Institutional KG
Entity relationships Decision rationales
Current state Trade-offs considered
Facts Why options were NOT chosen
What we know How we reason

Query examples: - "How would our past selves think about this problem?" - "What mistakes did we almost make in similar situations?" - "What assumptions have we invalidated recently?"

Current MIZOKI Implementation

Status: ✅ KG Brain Implemented, ⚠️ Reasoning History Not Captured

# knowledge_graph_brain_integration.py - Current Implementation
class KGNode:
    """Basic node with embeddings and properties"""
    node_id: str
    node_type: str
    embeddings: Optional[List[float]]
    properties: Dict[str, Any]

class KGEdge:
    """Edge with confidence and provenance"""
    edge_id: str
    source_id: str
    target_id: str
    edge_type: str
    confidence: float
    provenance: Optional[str]

Gap: Captures facts and relationships, but not reasoning history, rejected alternatives, or decision context.

Proposed Enhancement: Institutional Memory KG

@dataclass
class ReasoningHistoryNode(KGNode):
    """Captures HOW the organization reasons, not just WHAT it knows"""

    # Decision context
    decision_id: str                      # Link to original decision
    decision_timestamp: datetime
    decision_context: Dict[str, Any]      # Full context at decision time

    # Options considered
    options_evaluated: List[DecisionOption]
    selected_option: str
    selection_rationale: str

    # Rejected alternatives
    rejected_options: List[RejectedOption]
    rejection_reasons: Dict[str, str]

    # Confidence evolution
    initial_confidence: float
    final_confidence: float
    confidence_factors: List[str]

    # Meta-information
    reasoning_paradigm: str               # CoT, ToT, GoT, etc.
    key_assumptions: List[str]
    assumption_validations: Dict[str, bool]


@dataclass
class RejectedOption:
    """Records why an option was NOT chosen"""
    option_id: str
    option_description: str
    rejection_reason: str
    rejection_confidence: float
    conditions_for_reconsideration: Optional[str]  # When might this become viable?


@dataclass
class InstitutionalMemoryEdge(KGEdge):
    """Edge that captures reasoning relationship"""

    # Reasoning metadata
    discovered_how: str                   # How was this relationship discovered?
    discovery_confidence: float           # How confident at discovery?
    current_confidence: float             # Current confidence after validation
    validation_history: List[ValidationEvent]

    # Counterfactual information
    if_not_true: str                      # What would be different if this edge were false?
    invalidation_conditions: List[str]   # Conditions that would invalidate this

    # Usage tracking
    times_used_in_reasoning: int
    times_validated_by_outcome: int
    times_contradicted_by_outcome: int


class InstitutionalMemoryKG:
    """KG that captures organizational reasoning, not just facts"""

    def record_decision(
        self,
        decision: Decision,
        context: DecisionContext,
        options: List[DecisionOption],
        selected: str,
        rejected: List[RejectedOption]
    ) -> ReasoningHistoryNode:
        """Record a decision with full reasoning context"""

        node = ReasoningHistoryNode(
            node_id=f"decision:{decision.id}",
            node_type="reasoning_history",
            decision_id=decision.id,
            decision_timestamp=datetime.utcnow(),
            decision_context=context.to_dict(),
            options_evaluated=options,
            selected_option=selected,
            selection_rationale=decision.rationale,
            rejected_options=rejected,
            rejection_reasons={r.option_id: r.rejection_reason for r in rejected},
            initial_confidence=decision.initial_confidence,
            final_confidence=decision.final_confidence,
            reasoning_paradigm=context.paradigm,
            key_assumptions=decision.assumptions
        )

        self.add_node(node)

        # Link to relevant entity nodes
        for entity_id in decision.affected_entities:
            self.add_edge(InstitutionalMemoryEdge(
                source_id=node.node_id,
                target_id=entity_id,
                edge_type="DECIDED_FOR",
                discovered_how="decision_recording"
            ))

        return node

    def query_past_reasoning(
        self,
        problem_description: str,
        similarity_threshold: float = 0.7
    ) -> List[ReasoningHistoryNode]:
        """
        Query: "How would our past selves think about this problem?"

        Returns similar past decisions with full reasoning context.
        """

        # Embed the problem
        problem_embedding = self.embedding_service.embed(problem_description)

        # Find similar reasoning history nodes
        similar_nodes = []
        for node in self.get_nodes_by_type("reasoning_history"):
            context_embedding = self.embedding_service.embed(
                str(node.decision_context)
            )
            similarity = cosine_similarity(problem_embedding, context_embedding)

            if similarity >= similarity_threshold:
                similar_nodes.append((node, similarity))

        # Sort by similarity
        similar_nodes.sort(key=lambda x: x[1], reverse=True)

        return [node for node, _ in similar_nodes[:10]]

    def query_rejected_options(
        self,
        current_context: DecisionContext
    ) -> List[RejectedOption]:
        """
        Query: "What options have we rejected in similar situations?"

        Returns rejected options that might be worth reconsidering.
        """

        past_decisions = self.query_past_reasoning(str(current_context))

        reconsidered = []
        for decision in past_decisions:
            for rejected in decision.rejected_options:
                # Check if conditions for reconsideration are now met
                if rejected.conditions_for_reconsideration:
                    if self._check_conditions_met(
                        rejected.conditions_for_reconsideration,
                        current_context
                    ):
                        reconsidered.append(rejected)

        return reconsidered

    def query_invalidated_assumptions(
        self,
        domain: str,
        lookback_days: int = 30
    ) -> List[Tuple[str, str]]:
        """
        Query: "What assumptions have we invalidated recently?"

        Returns (assumption, invalidation_reason) pairs.
        """

        cutoff = datetime.utcnow() - timedelta(days=lookback_days)

        invalidated = []
        for node in self.get_nodes_by_type("reasoning_history"):
            if node.decision_timestamp < cutoff:
                continue

            for assumption, validated in node.assumption_validations.items():
                if not validated:
                    invalidated.append((
                        assumption,
                        f"Invalidated in decision {node.decision_id}"
                    ))

        return invalidated

6. Pattern 5: Continuous Improvement via Meta-Learning Loops

Concept

Systems improve not just outcomes, but:

Level What Improves
Object-level Individual decisions
Process-level How decisions are made
Meta-level How the process improves

Meta-learning loops periodically audit: - Research loop structure - Evidence standards - Coordination patterns - Evaluation thresholds

Current MIZOKI Implementation

Status: ✅ Dual-horizon Cycles, ⚠️ No Process Improvement Loop

# mizoki_learning_stability.py - Current Implementation
class DualHorizonCycles:
    """Tactical (24h) + Strategic (7d) cycles"""

    tactical_metrics = ['pain_points_addressed', 'mitigations_proposed']
    strategic_metrics = ['kg_refactoring', 'hypothesis_invalidation']

Gap: Cycles improve outcomes, but don't improve the research process itself.

Proposed Enhancement: Meta-Learning Controller

@dataclass
class ProcessMetrics:
    """Metrics about the research process, not outcomes"""

    # Research efficiency
    avg_time_to_hypothesis: timedelta
    avg_hypotheses_per_pain: float
    hypothesis_validation_rate: float

    # Coordination quality
    role_handoff_smoothness: float       # 0-1, higher = cleaner handoffs
    agent_utilization: float             # % of agent capacity used
    coordination_overhead: float          # Time spent coordinating vs researching

    # Evidence quality
    avg_evidence_depth: float            # Layers of supporting evidence
    evidence_contradiction_rate: float   # % of evidence later contradicted
    evidence_reuse_rate: float           # % of evidence reused in other research

    # Evaluation calibration
    prediction_accuracy: float           # How well predictions match outcomes
    confidence_calibration: float        # Is stated confidence accurate?
    escalation_appropriateness: float    # Were escalations necessary/sufficient?


class MetaLearningController:
    """Improves the research process itself, not just outcomes"""

    audit_interval_days: int = 7
    adjustment_rate: float = 0.05        # 5% max change per audit
    min_samples_for_adjustment: int = 20

    def audit_process(
        self,
        period_start: datetime,
        period_end: datetime
    ) -> ProcessAuditResult:
        """Audit research process for improvement opportunities"""

        # Collect process metrics
        metrics = self._compute_process_metrics(period_start, period_end)

        # Compare to baselines
        baselines = self._get_baselines()
        deviations = self._compute_deviations(metrics, baselines)

        # Identify improvement opportunities
        opportunities = []

        # Research efficiency
        if metrics.avg_hypotheses_per_pain > 5:
            opportunities.append(ProcessImprovement(
                area="hypothesis_generation",
                issue="Too many hypotheses generated",
                recommendation="Tighten initial filtering criteria",
                adjustment={"hypothesis_confidence_threshold": +0.05}
            ))

        # Coordination quality
        if metrics.coordination_overhead > 0.3:
            opportunities.append(ProcessImprovement(
                area="coordination",
                issue="High coordination overhead",
                recommendation="Reduce handoff frequency, increase agent autonomy",
                adjustment={"autonomy_threshold": -0.05}
            ))

        # Evidence quality
        if metrics.evidence_contradiction_rate > 0.2:
            opportunities.append(ProcessImprovement(
                area="evidence",
                issue="High evidence contradiction rate",
                recommendation="Increase evidence validation before use",
                adjustment={"evidence_validation_depth": +1}
            ))

        # Evaluation calibration
        if metrics.confidence_calibration < 0.7:
            opportunities.append(ProcessImprovement(
                area="evaluation",
                issue="Poor confidence calibration",
                recommendation="Adjust confidence scoring formula",
                adjustment={"confidence_dampening_factor": +0.1}
            ))

        return ProcessAuditResult(
            period=(period_start, period_end),
            metrics=metrics,
            deviations=deviations,
            improvements=opportunities
        )

    def apply_process_improvements(
        self,
        audit_result: ProcessAuditResult
    ) -> List[ProcessAdjustment]:
        """Apply approved process improvements"""

        adjustments = []

        for improvement in audit_result.improvements:
            # Check if we have enough samples
            samples = self._count_relevant_samples(improvement.area)
            if samples < self.min_samples_for_adjustment:
                continue

            # Apply adjustment with rate limiting
            for param, delta in improvement.adjustment.items():
                current = self.get_param(param)
                new_value = current + delta * self.adjustment_rate

                # Validate bounds
                new_value = self._clamp_to_bounds(param, new_value)

                self.set_param(param, new_value)
                adjustments.append(ProcessAdjustment(
                    param=param,
                    old_value=current,
                    new_value=new_value,
                    reason=improvement.recommendation
                ))

        return adjustments

    def audit_meta_process(self) -> MetaAuditResult:
        """Audit the auditing process itself (meta-meta-learning)"""

        # How effective have our process improvements been?
        recent_audits = self._get_recent_audits(n=10)

        improvement_effectiveness = []
        for audit in recent_audits:
            for adjustment in audit.applied_adjustments:
                # Compare metrics before and after adjustment
                before = self._get_metrics_at(audit.period[0])
                after = self._get_metrics_at(audit.period[1])

                effectiveness = self._compute_adjustment_effectiveness(
                    adjustment, before, after
                )
                improvement_effectiveness.append(effectiveness)

        # Meta-recommendation
        avg_effectiveness = sum(improvement_effectiveness) / len(improvement_effectiveness)

        if avg_effectiveness < 0.3:
            meta_recommendation = "Process improvements not effective; consider fundamental restructure"
        elif avg_effectiveness < 0.6:
            meta_recommendation = "Process improvements partially effective; tune adjustment rate"
        else:
            meta_recommendation = "Process improvements effective; maintain current approach"

        return MetaAuditResult(
            improvement_effectiveness=avg_effectiveness,
            recommendation=meta_recommendation,
            suggested_adjustment_rate=self._suggest_adjustment_rate(avg_effectiveness)
        )

7. Implementation Roadmap for MIZOKI

Phase 1: Close Learning Feedback Loop (Week 1-2)

Objective: Wire learning signals back into decision-making

# Integration point in research_ops_contract.py
class EnhancedResearchOrchestrator:

    def __init__(self):
        self.learning_feedback = LearningFeedbackLoop()

    async def process_pain_point(self, pain: NormalizedPainPoint):
        # Enrich with curriculum metadata
        curriculum_pain = self.learning_feedback.enrich_with_curriculum(pain)

        # Prioritize with learning yield
        priority = curriculum_pain.curriculum_priority(
            self.learning_feedback.get_agent_competence()
        )

        # Continue with existing flow...

Deliverables: - [ ] LearningFeedbackLoop class - [ ] Integration with PainPointNormalizer - [ ] Integration with AgentMarketOrchestrator - [ ] Firestore schema for competence tracking

Phase 2: Institutional Memory KG (Week 3-4)

Objective: Capture reasoning history, not just facts

# Integration in knowledge_graph_brain_integration.py
class EnhancedKGBrain:

    def record_decision_with_history(self, decision: Decision):
        # Existing: Record decision
        self.record_decision(decision)

        # New: Record reasoning history
        self.institutional_memory.record_decision(
            decision=decision,
            context=self.current_context,
            options=decision.options_evaluated,
            selected=decision.selected,
            rejected=decision.rejected_options
        )

Deliverables: - [ ] ReasoningHistoryNode schema - [ ] InstitutionalMemoryKG class - [ ] Query methods for past reasoning - [ ] Migration script for existing decisions

Phase 3: Learning Yield Evaluation (Week 5-6)

Objective: Measure knowledge gained, not just outcomes

# Integration in autonomous_research_patterns_v2.py
class EnhancedResearchPipeline:

    async def evaluate_complete(self, outcome: MitigationOutcome):
        # Existing: Task evaluation
        task_score = await self.evaluate_task(outcome)

        # New: Learning yield evaluation
        learning_yield = await self.learning_evaluator.evaluate_research_outcome(
            research_session=self.current_session,
            outcome=outcome,
            kg_before=self.kg_snapshot_before,
            kg_after=self.kg_snapshot_after
        )

        # Composite score
        return self.institutional_evaluator.evaluate_complete(
            outcome=outcome,
            learning_yield=learning_yield
        )

Deliverables: - [ ] LearningYieldMetrics class - [ ] LearningYieldEvaluator class - [ ] InstitutionalEvaluator class - [ ] Dashboard for learning yield visualization

Phase 4: Meta-Learning Controller (Week 7-8)

Objective: Improve the research process itself

# New module: meta_learning_controller.py
class MetaLearningOrchestrator:

    def __init__(self):
        self.controller = MetaLearningController()

    async def run_weekly_audit(self):
        # Audit process metrics
        audit = await self.controller.audit_process(
            period_start=datetime.utcnow() - timedelta(days=7),
            period_end=datetime.utcnow()
        )

        # Apply improvements
        adjustments = await self.controller.apply_process_improvements(audit)

        # Log for visibility
        await self.log_audit_results(audit, adjustments)

        return audit

Deliverables: - [ ] MetaLearningController class - [ ] Process metrics collection - [ ] Weekly audit job - [ ] Process improvement application

Phase 5: Role Rotation Policy (Week 9-10)

Objective: Build skills through deliberate rotation

# Integration in mizoki_learning_stability.py
class EnhancedAgentMarketOrchestrator:

    def __init__(self):
        self.rotation_policy = RoleRotationPolicy()

    def assign_role(self, agent_id: str, task: ResearchTask):
        # Get agent history
        history = self.get_agent_history(agent_id)

        # Select role with rotation consideration
        role = self.rotation_policy.select_role(
            agent=history,
            task=task,
            available_roles=self.available_roles
        )

        return role

Deliverables: - [ ] RoleRotationPolicy class - [ ] AgentRoleHistory tracking - [ ] Skill decay mechanism - [ ] Rotation debt tracking


8. Success Metrics

Institutional Learning KPIs

Metric Baseline Target Measurement
Learning Yield per Research N/A 0.5+ Average LearningYieldMetrics.compute_learning_yield()
Knowledge Reuse Rate N/A 30%+ % of insights reused in subsequent research
Agent Skill Breadth 3 roles 6+ roles Avg roles per agent with skill > 0.5
Reasoning History Coverage 0% 80%+ % of decisions with full reasoning context
Process Improvement Effectiveness N/A 60%+ % of process changes that improve metrics
Curriculum Completion N/A 70%+ % of skill categories exercised per quarter

Compound Advantage Indicators

Indicator Description Target Trend
Time to Resolution Avg time to resolve similar pain types ↓ 20% quarterly
First-attempt Success % of mitigations that work on first try ↑ 10% quarterly
Novel Pain Handling Success rate on never-before-seen problems ↑ 5% quarterly
Cross-Domain Transfer Success when applying lessons from other domains ↑ 15% quarterly

9. Conclusion

The patterns documented here represent a paradigm shift from autonomous agents as tools to autonomous agents as institutions. By implementing:

  1. Pain as Curriculum - Every operation contributes to capability building
  2. Role Rotation - Agents develop cross-functional judgment
  3. Learning Yield Evaluation - Success measured by knowledge gained
  4. Institutional KG - Reasoning history persisted for future reference
  5. Meta-Learning - The process improves itself continuously

MIZOKI can evolve from a sophisticated automation platform into a living organizational mind that remembers, reasons, adapts, and improves continuously—even as people, systems, and priorities change.

Strategic Outcome: Compound organizational intelligence that outpaces competitors through accumulated judgment, not just accumulated data.


References

  1. MIZOKI Codebase Analysis (January 2026)
  2. Autonomous Research Patterns V2 (autonomous_research_patterns_v2.py)
  3. Learning Stability Framework (mizoki_learning_stability.py)
  4. Research Ops Contract (research_ops_contract.py)
  5. Knowledge Graph Brain Integration (knowledge_graph_brain_integration.py)
← All docsView source on GitHub →