MIZOKI audit remediation — execution prompts
8 paste-ready prompts. 4 for Claude Code (CC-1..CC-4), 4 for Codex (CX-1..CX-4). Together they cover all 38 work orders in docs/audits/AUDIT_WORK_ORDERS_2026-09-08.md except WO-40..43 (customer-loop proof — human/owner work, not agent work).
Prepared 8 Sep 2026 for Boss. Repo: MIZOKI-3-5/MIZOKICloudRun. Source audit: docs/audits/ + project doc claude/mizoki-repo-audit-profit-roadmap-2026-09-06.md.
Split and run order
| Prompt | Tool | Package | Work orders | Owns these paths (no other prompt touches them) |
|---|---|---|---|---|
| CX-1 | Codex | A — truth | WO-00, 30, 31 | docs/audits/**, OPEN_ITEMS.md, docs/BUILD_DEBT.md, miz-oki-command-center-ui/app/api/bff/lanes/** |
| CC-1 | Claude Code | B — spending gates | WO-01..07 | services/service-policy-engine/**, services/service-validation-orchestrator/**, services/service-decision-control-plane/**, services/service-approval-routing/**, services/service-action-runner/main.py |
| CC-2 | Claude Code | C — tenant + durable execution (backend) | WO-17, 18, 19, 20, 32 | services/service-action-runner/execution_adapters/{base,portfolio,meta_ads,registry,flags}.py, shared tenant resolver, services/service-policy-engine/pacing_veto.py |
| CX-2 | Codex | C — UI authz + F product truth | WO-15, 16, 22, 23, 33 | miz-oki-command-center-ui/app/api/** (except bff/lanes/**), website pilot intake |
| CC-3 | Claude Code | D — economics | WO-08, 09 | services/net-yield/** |
| CC-4 | Claude Code | E — causal | WO-10..14, 29 | services/measurement-rails/**, services/service-media-incrementality/**, cell 26/27 modules, src/shared/growth_control/f2_ltv/**, services/lift-engine/** |
| CX-3 | Codex | C/F — perimeter + providers | WO-21, 24, 25 | miz-oki-adk-agents/boss/**, miz-oki-adk-agents/app/main.py, services/service-action-runner/execution_adapters/{google_ads,credentials}.py, services/service-data-manager-connector/**, miz-oki-adk-agents/kg-canonical-ingest/mappers/google_ads.py |
| CX-4 | Codex | H — release gate | WO-26, 27, 28 | .github/workflows/**, tests/test_client_library_pin_ratchet.py, per-service requirements*.txt lockfiles |
Order: run CX-1 first (it re-anchors every finding to current main and writes the reconciliation table the others read). Then CC-1 + CC-2 + CX-2 + CX-3 in parallel (disjoint paths). Then CC-3 + CC-4 + CX-4 in parallel. CC-1 and CC-2 share services/service-action-runner/ but not files — CC-1 owns main.py, CC-2 owns execution_adapters/. CX-3 also touches execution_adapters/google_ads.py and credentials.py only.
Merge policy for every prompt: work on a audit/<prompt-id>-<slug> branch (never claude/* — those auto-merge to main within seconds). Open a PR. Do not merge. Boss merges after APPROVED: [MERGE] (typed by the owner without the brackets; the bracketed form keeps the gate-leak ratchet honest). Anything under .github/**, deployment/terraform/**, deployment/cloudbuild*, CODEOWNERS is a protected path and always goes through a review PR. Merge ≠ deploy; deploys are dispatch-only and owner-triggered.
COMMON HEADER — paste at the top of every prompt
You are working in the MIZOKI-3-5/MIZOKICloudRun repository (MIZ OKI 3.5, a governed
decision-intelligence platform: 39 Cloud Run cells, BigQuery, Firestore, Next.js command-center UI).
Read in this order before touching code (AGENTS.md 1.1 authority order): CONSTITUTION.md (Article VI
governs governance surfaces), AGENTS.md, OPERATING_SYSTEM.md, GOVERNANCE.md, TRUTH.md,
.claude/skills/miz-oki-platform-expert/SKILL.md, README.md, CLAUDE.md + CLAUDE_MEMORY.md and every
file the memory router returns; then docs/audits/AUDIT_WORK_ORDERS_2026-09-08.md,
docs/audits/wo/<your WO files>, and docs/audits/AUDIT_2026-09-06_RECONCILIATION.md (CX-1 wrote it).
Ground rules
- Audit findings are HYPOTHESES pinned at commit fc8b03f9. main has moved. Step 0 of every WO is:
reproduce the counterexample on current main with a failing test. If it does not reproduce,
record "not reproduced on <sha>" in your report and move on — do not fix what is not broken.
- Fail closed. Every fix must make a refusal path explicit and tested. Never widen access to make a test pass.
- No new architecture. Reuse the existing modules named in the WO. If a WO says "reuse X", reuse X.
- Tests are the deliverable. Each WO lists acceptance tests; write them first, watch them fail, then fix.
Name them test_wo<nn>_<what>. Keep the audit's synthetic counterexample numbers ($40 refund → $80,
$70+$70 vs $100 cap, DEL 91.7, etc.) as fixtures so the regression is recognizable.
- Branch: audit/<PROMPT-ID>-<slug>. NEVER use a claude/* branch (they auto-merge to main in seconds).
Commit per WO with message "WO-nn: <title>". Open ONE PR for the prompt when done. Do not merge.
- Protected paths (.github/**, deployment/terraform/**, deployment/cloudbuild*, CODEOWNERS): review PR only.
- Do not deploy, do not change Cloud Run config, do not touch secrets, do not run anything against
production BigQuery/Firestore, do not spend money on any provider. Local + test fixtures only.
- Coordination: before starting, run
python scripts/claude_memory.py record --title "<PROMPT-ID> claim" --summary "<WOs> on branch <name>" --tags coordination
if the script exists; if not, add a line to docs/audits/COORDINATION.md.
- Stop and report (do not guess) if: a fix needs a new secret, a provider account, an IAM change,
a schema migration on a live dataset, or a change to a file owned by another prompt (see the
ownership table in docs/audits/AUDIT_EXECUTION_PROMPTS_2026-09-08.md).
Final report — write docs/audits/reports/<PROMPT-ID>_REPORT_<date>.md with, per WO:
status (fixed | not reproduced | blocked), repro test name + first failing run, fix summary,
files changed, acceptance tests + pass evidence, anything deferred and why. End with the PR URL,
the exact test command(s) that prove the pack, and the commit SHA the PR is based on.
CC-1 · Claude Code · Package B — spending admission hard gates (WO-01..07)
<COMMON HEADER>
PROMPT-ID: CC-1. Branch: audit/cc-1-spending-gates.
Work orders: WO-01, WO-02, WO-03, WO-04, WO-05, WO-06, WO-07. All P0. Lane: ENG + SEC.
You own: services/service-policy-engine/**, services/service-validation-orchestrator/**,
services/service-decision-control-plane/**, services/service-approval-routing/**,
services/service-action-runner/main.py, tests/governance/test_decision_control_plane.py,
tests/governance/test_action_runner.py. Do not edit services/service-action-runner/execution_adapters/**
(CC-2/CX-3 own it) — if a fix needs it, write the interface you need in main.py and report the gap.
Why this pack exists: the audit showed that a proposal can be authorized and spend money while a hard
economic/consent/identity check has failed, with unbound evidence, an asserted approver, and a
double-redeemed approval. This pack makes every one of those a terminal refusal.
WO-01 Hard-gate failures are terminal
Files: services/service-policy-engine/main.py (pass-rate / DEL eligibility),
services/service-validation-orchestrator/main.py (six-check media validator; incremental_profit check).
Step 0: build a fixture where the incremental-profit check FAILS and the other five checks pass at
100%; assert policy currently returns ELIGIBLE (audit: DEL 91.7). That is the failing test.
Fix: introduce an explicit check taxonomy: HARD = {economic (incremental profit, treasury), integrity,
consent, policy}, RANK = everything else. Compute eligibility as: if any HARD check failed →
INELIGIBLE with reason codes, before DEL is computed; DEL only ranks candidates that passed.
Do not implement this as a weight tweak. The validator must emit per-check {name, class, passed}.
Accept: (a) each HARD check failing alone, all others 100% → INELIGIBLE; (b) hypothesis/property test:
for any vector of RANK scores, a HARD failure never flips to ELIGIBLE; (c) existing eligible
fixtures still pass (no regression in the happy path).
WO-02 Bounded exploration class
Files: services/service-policy-engine/main.py.
Fix: eligibility_class ∈ {standard, exploration}. exploration requires envelope_id (approved
exploration budget record), cap, and a logged assignment probability; it is NOT exempt from HARD
checks. Store the class on the decision record.
Accept: exploration candidate without envelope_id → refused; with envelope over cap → refused;
HARD failure under exploration → refused.
WO-03 Bind evidence passports to the decision
Files: services/service-decision-control-plane/main.py, decision_meter.py.
Step 0: reproduce: supply a passport for a different tenant with an invalid seal and stage the actuator
registry so Stage-4 is reachable; assert a signed authorization is currently issued.
Fix: resolve the passport as an immutable record by id; verify seal; require passport.tenant ==
decision.tenant, passport.action_fingerprint == fingerprint(decision.action), passport.model_version
and horizon present, validity window covers now. Put those bound fields INSIDE the signed
authorization payload so a later reader can re-verify. Distinct reason codes:
PASSPORT_FOREIGN_TENANT, PASSPORT_SEAL_INVALID, PASSPORT_STALE, PASSPORT_ACTION_MISMATCH, PASSPORT_ALTERED.
Accept: one test per reason code → refused; a valid bound passport → authorized and the signature
verifies over the binding fields; tampering any bound field after signing → verification fails.
WO-04 Holdout registration is proved, not asserted
Files: services/service-decision-control-plane/main.py (experiment sufficiency), and the interface
services/service-action-runner/main.py uses to check holdouts. If the check lives in
execution_adapters/base.py, do NOT edit it — expose a resolver in main.py and report.
Fix: sufficiency = registry lookup of holdout_id returning {tenant, registered_at, salt_version}
with registered_at < first_exposure_at and tenant match. The proposer boolean is ignored.
Accept: unregistered id, post-exposure registration, other-tenant registration → refused;
properly registered → passes.
WO-05 Approver identity from authentication only
Files: services/service-approval-routing/main.py + the principal-auth module it imports.
Step 0: HTTP test: service principal S sends approval with body.actor = "some human"; assert it currently succeeds.
Fix: approver identity and role come from the verified principal only. Body actor fields are either
ignored or must equal the principal (mismatch → 400). A service principal can never satisfy a
HUMAN approval requirement; a human principal without the required role → 403.
Accept: three HTTP tests: service+body-human → 403; human-wrong-role → 403; human-right-role → 200
and the stored approval carries the principal's verified id, not the body string.
WO-06 Rollback proof artifact before promotion
Files: services/service-action-runner/main.py (promotion path), tests/governance/test_action_runner.py.
Fix: promotion to any executing stage requires a stored RollbackProof {drill_id, tenant, account,
action_class, executed_at, outcome=success, evidence_ref} matching the exact tenant/account/action
class. Registration's rollback_demonstrated flag becomes advisory metadata only.
Accept: flag=True + no proof → refused; proof for a different action_class → refused;
matching proof → promotion allowed. ops/remediation/live_proof.py may be READ for the
proof shape; do not modify it.
WO-07 Single-use approval under concurrency
Files: services/service-decision-control-plane/main.py (redemption), its Firestore/DB access layer.
Step 0: with a fake transactional store, redeem the same approval from two threads; assert two
distinct authorization ids are issued today.
Fix: atomic claim (transaction/conditional write) on the approval record; authorization_id =
deterministic hash(approval_id, decision_fingerprint, tenant); the stored authorization is
returned verbatim on any retry, including after a simulated crash between claim and persist
(two-phase: claim → persist → mark complete; a retry that finds claim-without-persist completes it).
Accept: 50-way concurrent redemption → exactly one authorization id and 49 identical replays;
crash-after-claim retry → same id; crash-after-persist retry → same id; contention at every
write boundary covered by a fault-injection test.
Gates before the PR: full pytest for the four services + tests/governance; skill_sync.py --audit if
present (mizoki_canon.py --check is a no-op — do not cite it as a gate); ruff/black if configured.
PR title: "Audit pack B — spending admission hard gates (WO-01..07)". Include the report path.
CC-2 · Claude Code · Package C backend — tenant boundaries and durable execution (WO-17, 18, 19, 20, 32)
<COMMON HEADER>
PROMPT-ID: CC-2. Branch: audit/cc-2-tenant-durable-execution.
Work orders: WO-17, WO-18, WO-19, WO-20, WO-32. All P0 except WO-32 (P2). Lane: ENG + SEC.
You own: services/service-action-runner/execution_adapters/{base,portfolio,meta_ads,registry,flags,
ratelimit,inventory_gate}.py, the shared tenant resolver (locate it: grep for "def resolve_tenant",
"TenantRegistry", "strict_mode" across src/, common/, services/service-decision-control-plane/
decision_meter.py, services/service-policy-engine/main.py — it may be duplicated; consolidate to one
importable module under common/ or src/shared/ and make the others import it), and
services/service-policy-engine/pacing_veto.py. Do NOT edit google_ads.py or credentials.py (CX-3) or
service-action-runner/main.py (CC-1). If DCP main.py must change for WO-18, make the smallest
possible change and flag it in the report — CC-1 is editing that file concurrently.
Why this pack exists: the audit showed exposure caps and freezes living in process memory (two $70
checks pass a $100 cap → $140 exposure), a tenant resolver that accepts any tenant when its registry
read fails, treasury checked from one source at policy time and another at execution time, and an
execute() exception path that skips the freeze handler.
WO-17 Fail closed on empty/unavailable tenant registry
Step 0: mock the registry read to raise; assert resolve_tenant("anything") currently succeeds.
Fix: distinguish RegistryUnavailable from RegistryEmpty. In strict mode (make strict the default
for all serving paths; allow non-strict only under an explicit env flag documented in the module
docstring) unknown tenant → refused; unavailable → refused unless a last-good cache entry exists
with age < TENANT_REGISTRY_CACHE_TTL_S (default 300) — and log that the cache was used.
Add ownership enforcement: DCP reads of stored decisions and runner execute/rollback/outcome
writes must check record.tenant == caller tenant.
Accept: raise → refused; empty+strict → refused; cache within TTL → allowed with audit log;
cache past TTL → refused; cross-tenant read of a stored decision → 404; cross-tenant
rollback/outcome write → 403.
WO-18 One versioned constraint state, Decide → settlement
Step 0: show that a proposal can pass the policy-engine treasury check while DCP/Act holds no
reservation (the startup global-file path is empty/optional).
Fix: a single ConstraintResolver(tenant) returning {version, currency, treasury_cap, exposure_cap,
horizon, fetched_at} sourced from the tenant onboarding vault (the same source policy uses).
Admission, reservation, execution and settlement all call it and record constraint_version on
the action. Execution refuses if constraint_version != the version the reservation was made
under, or if fetched_at is older than CONSTRAINT_MAX_AGE_S. Delete or hard-deprecate the
optional global startup file path (leave a loud error if the env var is still set).
Accept: policy pass + no reservation → execution refused; version drift → refused; stale → refused;
matching → allowed and all four stages log the same version.
WO-19 Persistent atomic reservations and freezes
Files: execution_adapters/portfolio.py, base.py.
Step 0: two threads reserve $70 each against a $100 cap using the current dict → both pass.
Fix: reservations and freezes move to a transactional store behind a small interface
(ReservationStore with reserve(tenant, account, amount, cap) → ok|refused atomically,
release(), freeze(tenant, account, reason), is_frozen()). Provide an in-memory transactional
fake for tests and a Firestore implementation (transactions/conditional writes) — do NOT run it
against a live project; unit-test the Firestore implementation with the emulator or a mock
that enforces transaction semantics.
Accept: 2 threads/2 processes $70+$70 vs $100 → exactly one passes; restart mid-reservation
(drop the process-local object, re-instantiate) → reservation still held; freeze set in
one instance is visible in another.
WO-20 Cover the whole mutation-and-verification interval
Files: execution_adapters/base.py (and meta_ads.py as the reference adapter).
Step 0: raise an ambiguous exception (e.g., timeout after the provider call) inside adapter.execute
and show the freeze handler is not entered.
Fix: one guarded span: dispatch → provider call → read-back → verify, with the action state machine
proposed → validated → authorized → dispatched → confirmed | uncertain | failed →
compensated | closed persisted at every transition. Any exception after dispatch → state
uncertain + freeze + reconcile-before-retry; retry is only permitted from a reconciled state.
Accept: fault-injection tests for: provider success + client timeout; crash after provider success;
duplicate delivery of the same authorization; each → no second provider mutation, freeze
recorded, state = uncertain until reconcile marks confirmed/failed.
WO-32 Certification evaluator enforced on every promotion
Files: execution_adapters/portfolio.py; the certification evaluator referenced from
decision_meter.py (read-only for you; if enforcement must live in DCP, write the call site
in portfolio/registry and report the DCP hook needed).
Fix: promotion calls the evaluator per tenant/account/action_class; no bypass path or flag.
Accept: promotion without a certification record → refused; with a record for another action_class → refused.
Gates: pytest services/service-action-runner services/service-policy-engine tests/governance;
skill_sync.py --audit if present. PR title: "Audit pack C (backend) — tenant boundaries + durable
execution (WO-17..20, 32)".
CC-3 · Claude Code · Package D — net-yield economics (WO-08, 09)
<COMMON HEADER>
PROMPT-ID: CC-3. Branch: audit/cc-3-net-yield-economics.
Work orders: WO-08, WO-09. Both P1. Lane: ENG (+ MEAS sign-off on invariants).
You own: services/net-yield/** only.
Why this pack exists: the audit inspected the generated SQL and the refund path and found tenant
pooling in rate aggregations and non-idempotent refund/fee/COGS writes. These corrupt the economics
every other lane trusts. Neither finding was executed against BigQuery — you will be the first to do
so, but ONLY against a test dataset you create with fixtures (never the production dataset).
WO-08 Tenant-key every aggregation
Files: compute.py, bq.py, test_compute.py.
Step 0: snapshot the generated SQL for the return-rate and order-rate stages; assert (failing) that
every GROUP BY and JOIN ON includes tenant_id. Also write the invariance test below and run it
against the BigQuery emulator or a scratch dataset (bq mk a dataset named
audit_wo08_<yourinitials>_<date>; delete it at the end).
Fix: carry tenant_id through every CTE, join, group-by and maturity window; the nightly tenant
parameter must scope the recomputed tables, not just the final MERGE.
Accept: (a) SQL snapshot test: tenant_id present in every GROUP BY/JOIN ON of every stage;
(b) two-tenant invariance: load tenants A and B; compute; replace ALL of B's rows with
different values; recompute; A's outputs are byte-identical. (c) existing 9 net-yield tests
still pass.
WO-09 Replay-safe refunds, fees and changed orders
Files: returns_adjustment.py, bq.py, cost_config.py, test_order_economics.py, test_bq.py.
Step 0 (reproduce all four, as failing tests): duplicate delivery of one $40 refund → $80;
refund arriving before its order → lost; two flat fee types → one overwrites the other;
order changed after COGS computed → stale COGS.
Fix: durable unique event ledger keyed by (tenant_id, provider, provider_event_id) with atomic
insert-if-absent; refund application is a deterministic materialization from the ledger, not an
increment; unmatched refunds are retained in a pending table and matched on order arrival;
order changes bump order_version and trigger a versioned recompute of COGS/fees/net; fees are
keyed by (tenant_id, fee_type) and summed, never overwritten.
Accept (against a real test DB or emulator): duplicate refund → exactly one ledger row, $40 applied
once; refund-before-order → matched when order lands; crash between ledger insert and
apply, then retry → idempotent; two fee types → both present; order change → COGS recomputed
and both versions retained.
Gates: pytest services/net-yield (all existing + new); SQL snapshot tests committed under
services/net-yield/tests/snapshots/. Drop the scratch dataset. PR title: "Audit pack D — tenant-safe,
replay-safe net-yield economics (WO-08, 09)". In the report, state explicitly whether each of the
four WO-09 counterexamples reproduced on current main.
CC-4 · Claude Code · Package E — causal and statistical consolidation (WO-10..14, 29)
<COMMON HEADER>
PROMPT-ID: CC-4. Branch: audit/cc-4-causal-consolidation.
Work orders: WO-10, WO-11, WO-12, WO-13, WO-14 (P1/P2), WO-29 (P3), WO-45 (P1, addendum 2026-09-08 —
docs/audits/wo/WO-45.md; retire the v22 uploadClickConversions rail behind the Data Manager path).
Lane: MEAS + ENG.
You own: services/measurement-rails/**, services/service-media-incrementality/**, the cell 26 and
cell 27 modules (locate via config/actual_urls.py and the cell registry; they may be in
src/cells/cell26, src/cells/cell27 or srpaldl-cells/), src/shared/growth_control/f2_ltv/**,
services/lift-engine/**, tests/governance/test_f2_ltv.py. Do not touch
services/service-validation-orchestrator/** (CC-1) — WO-11's export path there: read it, and if it
must change, write the change as a patch file under docs/audits/patches/ and report it.
Why this pack exists: the platform currently labels individual purchases "caused" vs "anticipated"
(reordering two timestamps moved $1,000 → $10), feeds that sum into Meridian as a calibration point,
evaluates a policy on its own training rows, and never refreshes the intervals promotion reads.
None of this is identifiable or honest. The fix is estimand discipline, not a new causal engine.
WO-10 Experiment-level estimand replaces individual labels
Files: measurement-rails/causal_credit.py, main.py, test_causal_credit.py.
Step 0: the permutation test — same cell, swap two purchase timestamps, assert summed "caused"
value changes (audit fixture: $1,000 → $10). Failing test.
Fix: compute incremental revenue/profit at the assignment-unit level: (treated mean − control mean)
× n_treated, with a seeded bootstrap or analytic CI; output schema {estimand:
"ATE_revenue"|"ATE_profit", point, ci_low, ci_high, n_treated, n_control, horizon, salt_version}.
Any per-purchase allocation that remains (for reporting) is emitted under
allocation_convention with label "convention" and must never feed downstream as causal.
Keep the 14 existing causal-credit tests passing or replace each with an explicit successor.
Accept: permutation invariance; schema fields present; convention output is structurally separate.
WO-11 Meridian calibration consumes effect + interval only
Files: service-media-incrementality/main.py (Meridian export).
Fix: export prior = {mean, sd} derived from WO-10's point and CI; refuse export when input is
labeled convention or lacks a CI; log the experiment id and n in the export record.
Accept: convention input → export refused with reason; valid experiment → prior populated, sd > 0.
WO-12 Cell 26: honest split and interval
Step 0: leakage test — assert evaluation rows ∩ training rows ≠ ∅ today.
Fix: grouped cross-fitting (groups = assignment unit); bootstrap over refits, not over predicted
individual effects; report exposure counts. Reuse the seeded-bootstrap and grouped
cross-fitting utilities in services/lift-engine — import them, do not copy.
Accept: leakage test passes (disjoint); interval widens monotonically as n shrinks in a synthetic
sweep; exposure counts present.
WO-13 Cell 27: intervals refresh on posterior update
Step 0: lifecycle test — after N updates with a clear winner, assert promotion still reads the
original interval and selects nothing.
Fix: recompute and persist intervals on every posterior update with a version; promotion reads
the latest version.
Accept: lifecycle test selects the winner; stale-version read is impossible (test asserts the
reader uses max version).
WO-14 F2 retention multiplier provenance
Files: src/shared/growth_control/f2_ltv/dtr.py, tests/governance/test_f2_ltv.py,
services/net-yield/test_returns_adjusted_f2_bridge.py (read-only; net-yield is CC-3's —
if the bridge test must change, report it).
Fix: each multiplier carries provenance ∈ {assumption, baseline, measured_effect} and a source ref;
proposal generation and bid-value paths accept measured_effect only; assumption/baseline are
scenario-only.
Accept: assumption-tagged multiplier reaching a bid-value path → refused; scenario path accepts all.
WO-29 Dosage estimator: nonlinear fixtures and support refusal
Files: services/lift-engine/continuous_dosage.py, src/core/continuous_dosage.py,
tests/test_continuous_dosage.py.
Fix: add saturation and carryover synthetic fixtures with known optima; refuse when the requested
dose is outside observed support; keep the constant-dose refusal.
Accept: saturation optimum recovered within 5%; out-of-support request → refused; existing 3.0
fixture still recovers ≈3.03.
Gates: pytest for every owned path; if a notebook-style cell has no test harness, add a minimal one.
PR title: "Audit pack E — estimand discipline and interval hygiene (WO-10..14, 29)". The report must
include a one-paragraph "estimand statement" for WO-10 that MEAS can sign.
CX-1 · Codex · Package A — establish current truth (WO-00, 30, 31)
<COMMON HEADER>
PROMPT-ID: CX-1. Branch: audit/cx-1-current-truth. RUN THIS FIRST — every other pack reads your output.
Work orders: WO-00 (P0), WO-30 (P2), WO-31 (P2). Lane: ENG + OWNER.
You own: docs/audits/** (new files only; do not edit AUDIT_WORK_ORDERS_2026-09-08.md or wo/*),
OPEN_ITEMS.md, docs/BUILD_DEBT.md, miz-oki-command-center-ui/app/api/bff/lanes/** and the
onboarding page's readiness display only.
WO-00 Pin the audit commit and diff to current main
1. git fetch; record main SHA. Confirm fc8b03f9da963f79d010df05fdb9e16c467e0d8f is an ancestor.
2. For each WO in docs/audits/wo/manifest.json, take the "Files" line, and produce
git diff --stat fc8b03f9..main -- <those paths>. Classify each WO: unchanged | moved (give new
path) | already fixed (cite the commit and the test that proves it) | file missing.
3. Write docs/audits/AUDIT_2026-09-06_RECONCILIATION.md: a table with one row per WO-nn and per
R0..R33: path at fc8b03f9, path at main, classification, evidence, owning prompt (from the
ownership table in docs/audits/AUDIT_EXECUTION_PROMPTS_2026-09-08.md). Also record: main SHA,
date, and the list of files that exist at main but not at fc8b03f9 in the owned paths.
4. Specifically resolve these unknown locations and write them into the table: the shared tenant
resolver (grep "def resolve_tenant", "TenantRegistry", "strict_mode"); cell 26 and cell 27
modules (config/actual_urls.py, cell registry, src/cells/, srpaldl-cells/); the site pilot
request code for WO-33 (not found under repo root — check the website repos referenced in
docs/ and CLAUDE.md, and record the repo+path or "not in this repo").
5. Do not create tickets; do not fix code.
Accept: all 38 WO rows and all 34 R rows filled; no "TBD".
WO-30 Lane-readiness evidence join
Files: miz-oki-command-center-ui/app/api/bff/lanes/status/route.ts (+ route.test.ts), onboarding page readiness display.
Step 0: Vitest: a tenant with no economics record; assert the lane currently reports "ready"
(or whatever the default is). If it already reports not_ready, record "not reproduced".
Fix: every readiness flag must be derived from a present, non-default evidence record with a
timestamp; missing/default → not_ready with a reason string shown in the UI.
Accept: no-economics tenant → not_ready; stale evidence (older than the lane's freshness window) →
not_ready:stale; full evidence → ready.
WO-31 Reconcile the open register
Files: OPEN_ITEMS.md, docs/BUILD_DEBT.md.
Fix: add one line per WO-nn under a new "Audit 2026-09-06 remediation" section linking to
docs/audits/wo/WO-nn.md and (once they exist) the GitHub issue numbers from
docs/audits/wo/issue_map.json; close/strike any register item superseded by the Sep 2 state
(e.g., #804 closed via #809/#810, F4 global params, supply-veto v2) with the superseding
reference. Do not create a second status document.
Accept: one register; every WO linked; no duplicate of an already-closed item; git diff shows only
additions and strike-throughs, no deletions of open items.
Gates: npx vitest run app/api/bff/lanes; markdown lint if configured.
PR title: "Audit pack A — reconciliation table, lane readiness, register (WO-00, 30, 31)".
CX-2 · Codex · UI authorization and product truth (WO-15, 16, 22, 23, 33)
<COMMON HEADER>
PROMPT-ID: CX-2. Branch: audit/cx-2-ui-authz-product-truth.
Work orders: WO-15 (P0), WO-16 (P0), WO-22 (P2), WO-23 (P2), WO-33 (P3). Lane: SEC + ENG + CUST.
You own: miz-oki-command-center-ui/app/api/** EXCEPT app/api/bff/lanes/** (CX-1), plus the website
pilot intake once CX-1's reconciliation table tells you where it lives. TypeScript/Next.js/Vitest.
Reference implementation for role + actor checks: app/api/bff/actions/authorize/route.ts and its
route.test.ts — copy its pattern, do not invent a new one.
WO-15 Role-check the onboarding mutation handlers
Files: app/api/bff/tenant-economics/save/route.ts; the tenant cost save route; the connector
credential save route under app/api/bff/connectors/ (find it: grep "credential" in
app/api/bff/connectors). Also the backend economics/cost routes these call — find them in
the BFF's fetch targets and note them; if they are Python services, write the required
backend change as a patch file under docs/audits/patches/ and report it (backend is CC/CX-3 territory).
Step 0: Vitest: construct a viewer identity for tenant T, stub the adapters, call all three save
handlers; assert 200 and that the adapter was invoked. Failing test.
Fix: apply the authorize route's role + actor + tenant check to all three; role must be admin (or
the role the authorize route uses for mutations); actor must be the verified principal.
Accept: viewer → 403 on all three and NO adapter call; admin → 200; wrong-tenant admin → 403.
WO-16 Route → required-role inventory enforced in CI
Fix: a script (scripts/ui_route_roles.ts or .py) that walks app/api/**/route.ts, detects mutating
handlers (POST/PUT/PATCH/DELETE), and requires each to export or annotate `requiredRole`
(pick the mechanism the authorize route already uses, or add a small `withRole()` wrapper);
emit a table to docs/audits/UI_ROUTE_ROLES.md. A Vitest test fails if any mutating route lacks it.
Accept: test green on your branch; the table lists every mutating route with a non-null role.
WO-22 Quarantine the legacy omnichannel allocator
Files: app/api/omnichannel/allocation/route.ts.
Step 0: call with no input; assert it currently returns a budget ($68,000) and projected revenue
($281,835.51); call with channel data missing clicks; assert a numeric conversion rate is
returned (267,000%).
Fix: missing live evidence → 422 {error: "evidence_unavailable", missing: [...]}; any simulated
output must carry {mode: "simulation"} and must never be returned from the live path; guard
every denominator (null, never a number); assert allocated + unallocated === budget or throw.
Accept: zero-input → 422; missing clicks → rate null; sum invariant test; simulation mode explicit.
WO-23 No path from the legacy allocator into DCP
Fix: a static test (Vitest or a small script in CI) asserting that no file outside
app/api/omnichannel/** imports the allocator's output type or calls its route, and that DCP
proposal intake never accepts an object with the allocator's shape. Add a deprecation banner
comment and a `X-MIZOKI-Legacy: allocation` response header.
Accept: CI test present and green; grep evidence in the report.
WO-33 Wire the site pilot request callback
Files: per CX-1's reconciliation table. If the code is in another repo, write the fix there on a
branch of the same name and report both PRs; if it truly does not exist, report "not in
scope" with evidence.
Fix: connect the request construction to the pilot intake route; on success persist a record;
on failure surface the error to the user (no silent drop).
Accept: submit test request → intake record exists; failure path shows an error.
Gates: npx vitest run; npx tsc --noEmit; next build if it runs in CI.
PR title: "Audit pack C/F (UI) — mutation role checks, allocator quarantine, pilot callback (WO-15,16,22,23,33)".
CX-3 · Codex · Boss perimeter and Google provider contracts (WO-21, 24, 25)
<COMMON HEADER>
PROMPT-ID: CX-3. Branch: audit/cx-3-perimeter-providers.
Work orders: WO-21 (P0), WO-24 (P2), WO-25 (P2). Lane: SEC + ENG + OPS.
You own: miz-oki-adk-agents/boss/**, miz-oki-adk-agents/app/main.py,
services/service-action-runner/execution_adapters/google_ads.py and credentials.py ONLY (CC-2 owns
the rest of execution_adapters/), services/service-data-manager-connector/**,
miz-oki-adk-agents/kg-canonical-ingest/mappers/google_ads.py. Deploy workflows are protected paths:
you may propose changes as a separate review PR but not in this branch.
WO-21 Authenticate the Boss service perimeter
Files: miz-oki-adk-agents/boss/app.py, miz-oki-adk-agents/app/main.py (the production entry that
re-exports the Boss app).
Step 0: FastAPI TestClient: POST the chat route and the direct-action route with no credentials;
assert 200 today.
Fix: (1) write docs/audits/BOSS_ROUTE_INVENTORY.md — every route, method, public|private,
required principal, and why; health/readiness are the only public routes unless the inventory
argues otherwise. (2) add an auth dependency (reuse the platform's existing principal-auth
module — find it via the DCP/approval services; do not write a new auth scheme) to chat and
direct-action and every route marked private. (3) In a SEPARATE review PR touching
.github/workflows/** and any deploy config, remove --allow-unauthenticated / allUsers for the
Boss service and document the invoker principal the UI uses. Do not merge; do not deploy.
Accept: anonymous chat → 401; anonymous direct-action → 401; authenticated principal → 200;
inventory committed; deploy-config PR opened with a one-paragraph rollback note.
Report: which routes changed from public to private, and the OPS verification step still needed
(read-only IAM invoker check on the live service — do not perform it).
WO-24 Google Ads adapter: leave sunset v21
Files: execution_adapters/google_ads.py (line ~55: API_VERSION default "v21"), credentials.py if
version-specific.
Step 0: assert the default is v21 (it is, as of main 15a6ba29b); check Google's sunset list.
Fix: default to the newest version that is BOTH listed in Google's current Ads API release notes
and supported by the installed google-ads client library (pin the library accordingly in the
runner's requirements — coordinate with CX-4's lockfile work by putting the pin in a clearly
marked single line). Audit every mutation and read-back call in the adapter against the new
version's changelog; fix renamed fields/enums. Keep GOOGLE_ADS_API_VERSION override.
Accept: unit tests for every mutation + read-back with recorded responses (VCR-style fixtures, no
live calls); a test asserting the default version is not in a committed SUNSET_VERSIONS
list; validate-only mutation path exists and is tested. Do NOT run against a real account;
write the validate-only run as an OPS checklist item in the report.
WO-25 Data Manager connector request contract
Files: services/service-data-manager-connector/main.py; kg-canonical-ingest/mappers/google_ads.py.
Step 0: unit test the built request against the Data Manager Event / IngestEvents / Destination REST
schema (write the schema as JSON-schema fixtures from the official docs): assert today it
omits productDestinationId, drops conversion_action, sends conversionValue as an object,
and does not encode hashed userData. Four failing tests.
Fix: conform exactly: productDestinationId per destination; conversion_action routed per event to
its destination (multi-action routing); conversionValue numeric with sibling currencyCode;
hashed userData normalized + SHA-256 + encoded as the schema requires; validate-only flag
plumbed and its diagnostics surfaced in the response.
Accept: four schema tests pass; a golden request fixture is committed; validate-only path tested
with a recorded diagnostics response. Do not build a second connector.
Gates: pytest for miz-oki-adk-agents/boss, the two adapter files, and the connector service.
PR title: "Audit pack C/F (perimeter + providers) — Boss auth, Ads API version, Data Manager contract
(WO-21, 24, 25)". Plus the separate protected-path PR for deploy config.
CX-4 · Codex · Package H — release gate (WO-26, 27, 28)
<COMMON HEADER>
PROMPT-ID: CX-4. Branch: audit/cx-4-release-gate. Everything here is a PROTECTED PATH → review PR only.
Work orders: WO-26 (P2), WO-27 (P2), WO-28 (P3). Lane: ENG + OPS.
You own: .github/workflows/**, tests/test_client_library_pin_ratchet.py,
tests/governance/test_deploy_allowlist_completeness.py (read for the pattern), per-service
requirements*.txt and lockfiles. CX-3 adds one pinned google-ads line in the runner's requirements —
merge around it, do not revert it.
WO-26 Protect the exact commit that ships
Files: .github/workflows/auto-merge-ai-branches.yml, auto-merge-grothendieck.yml,
auto-merge-mcclintock.yml, auto-sync-main.yml, ci.yaml.
Step 0: trace the push-retry path in each auto-merge workflow; document (in the report) the exact
step where a rebase can occur after the content gate ran. Write a workflow test (act, or a
scripted dry run) that demonstrates it.
Fix: after any rebase in the retry path, re-run the content gates on the new SHA or fail the job;
record the SHA that was gated in the job summary and assert it equals the pushed SHA.
Produce docs/audits/REQUIRED_CHECKS.md: the list of required status checks by name, and map
every application test suite (including the MCP/spec/origin suites the audit found unwired —
find them: grep for test directories not referenced by any workflow) to a CI job; wire the
unwired ones. Branch protection / rulesets cannot be set from the repo (plan-level 403) —
write the exact settings the owner must click as an OPS checklist in the report.
Accept: simulated rebase-after-check → job fails or re-gates; job summary shows gated SHA ==
pushed SHA; every test directory is referenced by at least one workflow (add a test that
asserts this).
WO-27 Dependency locking by service
Step 0: reproduce the count: scan non-archived requirements*.txt, count unbounded declarations
(no ==, no ~=, no upper bound); audit found 419 in 52 of 135 files.
Fix: for every service in the deploy allowlist (use test_deploy_allowlist_completeness.py's source
of truth), generate a lockfile (pip-compile or uv) from the existing requirements; extend
test_client_library_pin_ratchet.py to assert every allowlisted service has a lockfile and
that its Dockerfile installs from it; add a separate, non-blocking advisory-scan job (pip-audit)
that reports but does not gate yet.
Accept: ratchet test covers every allowlisted service; unbounded count in allowlisted services = 0;
non-allowlisted count reported, not gated; builds still succeed locally for at least three
services you pick (report which).
WO-28 Authenticated journey smoke per deploy
Files: ci.yaml + the Boss/UI deploy workflows.
Fix: a post-deploy job (workflow_dispatch + on deploy success) that runs an authenticated tenant
journey (login as the synthetic tenant → lane status → one read-only decision fetch) against
the named Cloud Run revision, using existing OIDC/WIF secrets — reference them by name, do not
create or print them. Output: revision name + pass/fail in the job summary.
Accept: workflow validates (actionlint); a dry-run mode that stubs the HTTP calls passes in CI;
the live mode is documented as owner-dispatched only.
Gates: actionlint on every changed workflow; pytest tests/test_client_library_pin_ratchet.py.
PR title: "Audit pack H — gate the shipped SHA, lock dependencies, journey smoke (WO-26, 27, 28)".
Because this is a protected path, the PR body must include the owner click-list for branch
protection and the rollback for each workflow change.
Not covered by any prompt (human work)
WO-40 select the design partner (CUST), WO-41 preregister the experiment (MEAS), WO-42 one complete decision-to-outcome trace (ENG, but only after every P0 in B and C is merged and green), WO-43 readout and promotion decision (OWNER). These start when CC-1, CC-2, CX-2 and CX-3 are merged.
After all eight PRs are open
- Merge order: CX-1 → CC-1, CC-2, CX-2, CX-3 (any order; resolve trivial conflicts in service-action-runner/ and DCP main.py) → CC-3, CC-4 → CX-4 last (it changes the gates the others were tested under).
- Re-run CX-1's reconciliation script against the merged main and update the table.
- Only then does WO-42 begin.
Addendum 2026-09-08 — WO-44..49 assignment (added after the review v1.1 source pass)
| Ticket | Goes to | Note |
|---|---|---|
| WO-44 | owner (rulings A/B, Team-plan upgrade, gcloud) → CX-4 for any workflow line |
Gates WO-26. Until ruled, homepage prod deploys are BLOCKED, not patched. |
| WO-45 | CC-4 (owns services/measurement-rails/**) |
Add to CC-4's list; the tree-wide sunset assertion lands with CX-3/WO-24. |
| WO-46 | owner/operator, before any cloud CC/CX run | App install on org MIZOKI-3-5; new session after. |
| WO-47 | CC-1 (owns services/service-decision-control-plane/**) |
Add to CC-1's list beside WO-03/WO-05. |
| WO-48 | CX-4 (owns .github/workflows/**) |
Same review PR as WO-26/27 where practical. |
| WO-49 | owner | Sequence AFTER all audit/* branches merge — history rewrite invalidates open branches. |