MIZ OKI 3.5 — Stage-2 domain benchmarks

Ref: MIZ-FIN-2026-005 Phase 13 · finance & CRE proof obligations.

This directory holds the benchmark artifacts that a human must accept before advisory_only can be lifted for a domain. Nothing in code lifts it — the conformance check pins advisory_only: true for finance and CRE (services/service-policy-engine/policy.yaml), and only a named approver accepting an artifact here that meets the bar changes that.

Why this exists (the gate)

Finance and CRE run their full validation batteries today, but purely advisory_only (Stage 3): the batteries exist, benchmark runs do not. The platform claim ceiling is built, pre-benchmark. Phase 13 supplies the harness (ops/remediation/benchmark_harness.py) that produces a labeled artifact; this README is the contract the artifact must satisfy and the labels it must carry.

The bar to lift advisory_only for a domain

An artifact is acceptance-eligible only if ALL hold, and even then the decision is a human's:

Gate Value (from policy.yaml)
DEL threshold finance/CRE ≥ 90
Min passport pass-rate floor ≥ 0.75
Sample size enough held-out windows to be meaningful (record N; a thin sample is a design target, not a benchmark result)
Point-in-time integrity every evidence read filtered ONLY on available_to_model_at — no look-ahead
Named approver recorded in the decision register (docs/INTEGRATION_PLAN.md §10)

DEL is 100 · (0.5·passport_pass_rate + 0.3·evidence_completeness + 0.2·verification_weight) — the same formula the policy engine scores live.

Claim labels (mandatory on every figure)

Every number an artifact reports carries exactly one of the five labels:

Artifact format

benchmark_harness.py writes one JSON file per run: docs/benchmarks/<domain>-<YYYYMMDD>-<runtag>.json with:

{
  "domain": "finance" | "cre",
  "run_tag": "...", "generated_at": "...",   // stamped by the caller, not the harness
  "label": "benchmark result" | "illustrative scenario" | "design target",
  "n_windows": <int>,
  "data_source": "<historical corpus id>" | "SYNTHETIC (no historical corpus)",
  "del": {"mean": .., "p05": .., "p50": .., "p95": .., "threshold": 90},
  "passport": {"pass_rate_mean": .., "floor": 0.75, "checks": [...]},
  "calibration": {...},          // predicted-vs-actual where outcomes exist
  "hit_rate": <float|null>,      // null when no ground-truth outcomes
  "acceptance_eligible": <bool>, // ALL gates above met AND label == benchmark result
  "reproduce": "python ops/remediation/benchmark_harness.py --domain <d> ..."
}

Data-gap status (2026-07-28)

At harness landing there is no historical finance/CRE outcome corpus in the repo to backtest against (registered here per the directive: "a thin benchmark honestly labeled beats a padded one"). Until a real corpus is wired, harness runs are labeled illustrative scenario — they prove the batteries score end-to-end through the governed path with point-in-time reads, and they do not support lifting advisory_only. See ops/remediation/benchmark_harness.py --help.

← All docsView source on GitHub →