SRPVDAL Evaluation Framework
Platform: MIZ OKI 3.5
Canonical loop: SENSE → REASON → PLAN → VALIDATE → DECIDE → ACT → LEARN
Rule: no validated evidence, no autonomous action.
1. Evaluation purpose
The evaluation framework prevents MIZ OKI 3.5 from confusing data availability with action readiness. A source can be connected and a plan can be plausible, but the system should not act unless the relevant gates pass.
Evaluation happens throughout the loop, with the highest concentration in VALIDATE and DECIDE.
2. Gate taxonomy
| Gate | Purpose | Applies to |
|---|---|---|
| Data quality | Confirm source data is complete, fresh, typed, deduplicated, and schema-compatible | SENSE |
| Connector health | Confirm auth, API version, quota, rate limits, pagination, retry behavior, and latency | SENSE |
| Canonicalization | Confirm raw records map to versioned canonical events with provenance | SENSE |
| Entity resolution | Confirm match keys, confidence, ambiguity, and merge behavior | SENSE, REASON |
| KG integrity | Confirm expected nodes, edges, evidence, and graph invariants | SENSE, REASON, LEARN |
| Reasoning quality | Confirm explanations have evidence and identify uncertainty | REASON |
| Planning quality | Confirm options are feasible, specific, reversible, and include a no-action baseline | PLAN |
| Policy validation | Confirm budget, brand, legal, privacy, safety, and approval rules | VALIDATE |
| Financial validation | Confirm spend, ROI, margin, risk, and downside constraints | VALIDATE |
| Statistical validation | Confirm confidence intervals, sample size, experiment quality, and regression risk | VALIDATE |
| Causal validation | Confirm uplift evidence, counterfactual support, OPE results, and causal confidence | VALIDATE |
| API/action validation | Confirm compatibility, idempotency, pre-state, post-state, and rollback plan | ACT |
| Outcome validation | Confirm predicted vs actual effect and detect drift or regression | LEARN |
3. Stage-by-stage requirements
SENSE evaluations
Required checks:
- source authentication status;
- source API version;
- schema compatibility;
- row/event counts;
- freshness;
- null and type errors;
- deduplication;
- raw payload hash or pointer;
- canonical event schema version;
- match-key completeness;
- consent/data-rights metadata when applicable.
Output artifact: source and canonicalization report.
REASON evaluations
Required checks:
- evidence path exists;
- evidence freshness;
- contradictory evidence;
- graph path confidence;
- causal vs correlational labeling;
- retrieval grounding quality;
- uncertainty statement.
Output artifact: reasoning evidence report.
PLAN evaluations
Required checks:
- candidate action specificity;
- objective alignment;
- no-action baseline;
- expected effect;
- required permissions;
- risk level;
- rollback or compensating action requirement.
Output artifact: candidate plan set.
VALIDATE evaluations
Required checks:
- policy pass/block/defer;
- budget and financial constraints;
- legal and brand constraints;
- safety and privacy constraints;
- statistical confidence;
- causal confidence;
- operational feasibility;
- required human approval.
Output artifact: validation decision record.
DECIDE evaluations
Required checks:
- selected candidate passed required gates;
- rejected alternatives are captured;
- confidence meets threshold;
- action class is allowed;
- approval state is clear;
- rationale is audit-ready.
Output artifact: decision receipt.
ACT evaluations
Required checks:
- pre-state captured;
- idempotency key present;
- operation compatibility;
- rate limit checked;
- approval token present where required;
- post-state captured;
- rollback handle or compensating action recorded.
Output artifact: action ledger record.
LEARN evaluations
Required checks:
- outcome observation window;
- predicted vs actual comparison;
- attribution or counterfactual baseline;
- calibration result;
- drift result;
- policy/model update recommendation;
- graph memory update.
Output artifact: learning record.
4. Minimum decision receipt
Every approved or blocked decision should have a receipt.
decision_trace_id: trace_...
stage: DECIDE
objective: maximize_incremental_profit_with_guardrails
selected_action:
type: budget_reallocation
target: campaign:123
payload_hash: sha256:...
status: approved_with_human_review
confidence: 0.82
expected_effect:
metric: incremental_roas
value: 2.1
validation:
data_quality: pass
policy: pass
financial: pass
causal: pass
legal_brand: pass
approval: pass
rejected_alternatives:
- type: pause_campaign
reason: insufficient causal confidence
approval:
required: true
approver_role: growth_owner
audit:
source_events: [...]
kg_evidence: [...]
created_at: "..."
5. Hard-block conditions
The system should block or defer action when:
- required source data is stale or missing;
- schema validation fails;
- source identity or account access is ambiguous;
- match confidence is below threshold for the target action;
- policy, legal, brand, privacy, or consent gates fail;
- financial downside exceeds limit;
- causal evidence is missing for a causal-impact claim;
- OPE or replay gate blocks promotion;
- approval is required but missing;
- pre-state cannot be captured;
- rollback or compensating action is unavailable for a high-risk operation.
6. Evaluation harness roadmap
Connector harness
- auth check;
- API version check;
- pagination test;
- rate-limit test;
- retry behavior;
- schema sample validation;
- raw payload preservation;
- canonical event count.
Canonical event harness
- schema validation;
- required fields;
- provenance fields;
- match keys;
- hash stability;
- immutable fact checks;
- version migration test.
KG harness
- expected node creation;
- expected edge creation;
- evidence attachment;
- confidence values;
- conflict behavior;
- duplicate behavior;
- stale edge detection.
Decision harness
- candidate generation;
- validation pass/block;
- approval path;
- decision receipt;
- action dry run;
- pre/post state;
- learning record.
Regression harness
- representative source fixture set;
- golden canonical event set;
- golden KG mapping set;
- expected gate result set;
- expected decision receipt set.
7. Reporting in the Command Center
The Command Center should expose:
- source health by cell;
- canonical event counts and failures;
- KG mapping status;
- validation gate matrix;
- pending approvals;
- blocked decisions and reasons;
- executed actions and rollback handles;
- predicted vs actual outcomes;
- learning updates;
- drift and regression alerts.
8. Operating principle
A fast decision is not valuable if it is not safe, explainable, reversible, and measurable. MIZ OKI 3.5 optimizes for governed decision velocity, not blind automation.