MIZ OKI 3.5 — Stage-2 domain benchmarks
Ref: MIZ-FIN-2026-005 Phase 13 · finance & CRE proof obligations.
This directory holds the benchmark artifacts that a human must accept before
advisory_only can be lifted for a domain. Nothing in code lifts it — the
conformance check pins advisory_only: true for finance and CRE
(services/service-policy-engine/policy.yaml), and only a named approver
accepting an artifact here that meets the bar changes that.
Why this exists (the gate)
Finance and CRE run their full validation batteries today, but purely
advisory_only (Stage 3): the batteries exist, benchmark runs do not. The
platform claim ceiling is built, pre-benchmark. Phase 13 supplies the
harness (ops/remediation/benchmark_harness.py) that produces a labeled
artifact; this README is the contract the artifact must satisfy and the labels
it must carry.
The bar to lift advisory_only for a domain
An artifact is acceptance-eligible only if ALL hold, and even then the decision is a human's:
| Gate | Value (from policy.yaml) |
|---|---|
| DEL threshold | finance/CRE ≥ 90 |
| Min passport pass-rate floor | ≥ 0.75 |
| Sample size | enough held-out windows to be meaningful (record N; a thin sample is a design target, not a benchmark result) |
| Point-in-time integrity | every evidence read filtered ONLY on available_to_model_at — no look-ahead |
| Named approver | recorded in the decision register (docs/INTEGRATION_PLAN.md §10) |
DEL is 100 · (0.5·passport_pass_rate + 0.3·evidence_completeness +
0.2·verification_weight) — the same formula the policy engine scores live.
Claim labels (mandatory on every figure)
Every number an artifact reports carries exactly one of the five labels:
- benchmark result — a real backtest over held-out historical windows with
a stated N. Only these can support lifting
advisory_only. - verified result — a value measured for one exact exercised path.
- pilot result — from a controlled pilot.
- design target — a goal, not a measurement (use when N is too thin to claim a benchmark).
- illustrative scenario — synthetic inputs on real infrastructure, proving the harness runs, not the domain's performance.
Artifact format
benchmark_harness.py writes one JSON file per run:
docs/benchmarks/<domain>-<YYYYMMDD>-<runtag>.json with:
{
"domain": "finance" | "cre",
"run_tag": "...", "generated_at": "...", // stamped by the caller, not the harness
"label": "benchmark result" | "illustrative scenario" | "design target",
"n_windows": <int>,
"data_source": "<historical corpus id>" | "SYNTHETIC (no historical corpus)",
"del": {"mean": .., "p05": .., "p50": .., "p95": .., "threshold": 90},
"passport": {"pass_rate_mean": .., "floor": 0.75, "checks": [...]},
"calibration": {...}, // predicted-vs-actual where outcomes exist
"hit_rate": <float|null>, // null when no ground-truth outcomes
"acceptance_eligible": <bool>, // ALL gates above met AND label == benchmark result
"reproduce": "python ops/remediation/benchmark_harness.py --domain <d> ..."
}
Data-gap status (2026-07-28)
At harness landing there is no historical finance/CRE outcome corpus in the
repo to backtest against (registered here per the directive: "a thin
benchmark honestly labeled beats a padded one"). Until a real corpus is wired,
harness runs are labeled illustrative scenario — they prove the batteries
score end-to-end through the governed path with point-in-time reads, and they
do not support lifting advisory_only. See
ops/remediation/benchmark_harness.py --help.