Completion Run — Phase A Matrix (Checkpoint A)
Run: completion-run (owner mission prompt, 2026-08-08)
Audited tree: main @ 6ef0275 (== origin/main at audit time; live mizoki3.com verified byte-identical to this tree for all 9 audited routes)
Auditors: three parallel read-only subagents — full reports in this directory (A1_site_content_audit.md, A2_platform_machinery_audit.md, A3_cross_reference_audit.md)
Governance basis: CONSTITUTION.md · TRUTH.md (claim labels; §5 enforcement; §6 language rules) · AGENTS.md Art. 3/7 · GOVERNANCE.md Art. 2/5 · .claude/rules/website-governance.md · story bank v1.1 (five rules) · net-yield positioning claim ledger
Verdict vocabulary: SHIPPED (exists + tests/evidence) · PARTIAL (code exists, mechanism incomplete/stub/untested) · MISSING (absent) · UNBACKED-CLAIM (live copy asserts machinery with no/insufficient backing) · DEMO-ONLY (backed solely by the site's deterministic demo runtime)
§1 — Corrections to the mission brief's assumed ground truth
The brief's "verified 2026-08-08" context was measured against the actual tree. Per CLAUDE.md §1 and TRUTH.md precedent 4.2/4.3, measurement wins:
| Brief assumed | Measured reality | Evidence |
|---|---|---|
scripts/content_qa.py MISSING |
SHIPPED at # MIZ OKI 3.5/scripts/content_qa.py (355 lines; rules A–D; self-test; gates deploy-homepage) |
A2 §a; run at HEAD: CONTENT QA OK — 12 scoped files clean, self-test PASS |
| All net-yield code MISSING | SHIPPED & COMPLETE per the B4 checklist: services/net-yield/ (7 endpoints, documented formula, strict cost-config validation, NET_YIELD_WRITEBACK=false + pinning tests), envelope order-economics block, BQ DDL, Shopify extender orders+refunds, docs/net-yield/{ADR-NY-001,RUNBOOK}.md |
A2 §d — "nothing missing" |
skills_sync.py unconfirmed |
SHIPPED; --check passes; SKILL.md v3.0 byte-identical across homes; docs/skills/BOSS_REGISTRATION.md exists |
A2 §b |
| Story bank / positioning docs to be committed from drop/ | Already canonical: drop/ copies byte-identical (md5-equal) to # MIZ OKI 3.5/docs/marketing/ copies; story bank carries v1.1 + all five rules |
A3 §1–2 |
| Briefing restructure to propose | Already applied — owner approved 2026-08-07; PROVE→PROFIT→ANTICIPATE woven into executive-briefing/js/data.js; canon re-pinned same commit |
A2 §e/f |
| Doorman blog draft to create | Published — blog/doorman-problem.html, posts.json 2026-08-03; doorman lines present on demo.html:152 and demo-signal.html:321 |
A2 §f |
| Shopify listing copy to create | SHIPPED — docs/marketing/shopify-app-listing-copy.md (content_qa-scoped) |
A2 §e |
| "Merging to main deploys production (site)" | FALSE for the site — deploy-homepage.yml is dispatch-only + typed APPROVED + canon/content_qa gates (since 2026-07-30). TRUE in a different way for branches: any claude/gemini/codex/copilot/cursor/** push auto-merges to main in seconds (-X theirs retry), and the Deploy Router then fires matching deploy-*.yml |
A2 §g; GOVERNANCE 2.1/2.2; AGENTS 3.1 |
| LII cells "pending" | Cells 33/34/35/36 exist with tests (61/60/46/27) and are live (dispatch-only deploys); cell28 remains legacy (never repurposed); mizoki-lii-mcp server absent (skills parity "yellow by design") |
A2 §c |
Still true from the brief: /signal/measurement makes present-tense machinery claims with insufficient backing (see §2) — this is the real build core. docs/BUILD_DEBT.md and docs/measurement-rails/ do not exist. MEASUREMENT_WRITEBACK flag does not exist anywhere.
§2 — Measurement-page machinery claims (the build-to-claim core)
Live /signal/measurement (byte-identical to repo) states at line 267 that "mechanisms on this page are operating machinery, and the windows and thresholds shown are operating defaults." Zero "Preview" strings on the page. Verdicts per claim (A2 Task 1, full evidence in A2 report):
| ID | Claim (page) | Verdict | Key evidence gap | Resolution (proposed) |
|---|---|---|---|---|
| M1 | House attribution recompute — 7d-click/1d-view, nightly, BQ unified.house_attribution, per-source lag profiles |
PARTIAL | Boss module's data fetch is an np.random mock; no 1-day-view logic; no nightly job; house_attribution table appears nowhere; lag-profiles JSON never loaded |
BUILD (B3.1) — real BQ-backed recompute job in new services/measurement-rails/, DDL for unified.house_attribution, per-source empirical lag profiles, nightly schedule via RUNBOOK (scheduler creation = operator) |
| M2 | Salted SHA-256 identity stitching; deterministic-only causal math; raw identifiers never persisted | BACKED | Real: identity_attribution cell (peppered SHA-256, privacy gate, 25 tests) + cell36 deterministic-only refusal (tested). Nuances: KMS pepper (not per-record salt); a parallel unsalted legacy stitcher in boss cross_platform_attribution_integration.py |
NO BUILD. Record boss legacy stitcher as BUILD_DEBT retirement item (supersede via rails util; migration-not-rename) |
| M3 | Drift monitor — platform vs model revenue, >20% for 3 consecutive days, alerting | PARTIAL (stub) | check_drift returns empty alerts + "future enhancement"; no 3-day logic; no Prometheus/Grafana drift wiring |
BUILD (B3.4) — daily comparison job, 3-consecutive-day alert artifact, Prometheus metric + Grafana alert-rule stub |
| M4 | Weight back-propagation of corrected conversion values to Meta/Google bidders | PARTIAL | Boss push_weights executor is log-only and NOT flag-gated; net-yield writeback is real+gated but net-yield-scoped, construction-only |
BUILD (B3.5) — rails module computing corrected values from house attribution + platform payload construction; hard MEASUREMENT_WRITEBACK=false default pinned by test; recommend/dry-run first; autonomy-ladder note. Boss log-only path recorded as BUILD_DEBT retirement item |
| M5 | Google Enhanced Conversions rail | PARTIAL | 1,712-line client (PIIHasher, OAuth, v22 adjustments, retries) exists; zero tests, no pipeline | BUILD (B3.3a) — rails google_enhanced_conversions.py (flag-gated, DRY_RUN default: build/validate/log, never send) reusing the client's payload spec; unit tests |
| M6 | Meta CAPI — shared event_id + 48h dedup | PARTIAL | event_id + Firestore dedup exist but coded window is 24h not 48h; second uploader omits event_id; no tests | BUILD (B3.3b) — rails meta_capi.py with shared event_id + 48h dedup, unit-tested; DRY_RUN default |
| M7 | AEM 8-event priority schema | PARTIAL | Priorities 1–8 + max-8 slice + LDU options coded in-module; no external config, no tests | BUILD (B3.3c) — externalized AEM priority-schema config + tests |
| M8 | GA4 Measurement Protocol, validated sends | PARTIAL | Three clients with debug/collect capability; none does validate-THEN-send; no tests | BUILD (B3.3d) — rails ga4_measurement_protocol.py with validate-before-send; tests |
| M9 | Offline conversions — GCLID/WBRAID/GBRAID, 90-day window | PARTIAL | ekis TS client real; boss Python uploader permanently mocked (client import commented out); no 90-day window anywhere | BUILD (B3.3e) — rails offline_conversions.py (three click-ID types, 90-day window enforcement, DRY_RUN default); tests |
| A10 | Designs registered before first impression (holdout / ghost bids / geo), MDE declared up front | PARTIAL | Holdouts real+tested (cell36 write-once, 10% permanent); geo/holdout/rct registry in service-media-incrementality; ghost bids = dataclasses only; MDE nowhere in code | BUILD-small (B3.6) — add mde declaration field to the experiment-registry design record + ghost_bid as a registrable design type (registration, not execution); ghost-bid execution path → BUILD_DEBT |
| A11 | Heterogeneous effects per segment via meta-learners | BACKED (via cell36) | cell36 estimators.py DRLearner+CUPED with tests; cell26 X/DR-Learner real but untested |
NO BUILD. cell26 test hardening → BUILD_DEBT |
| A12 | Automated refutation battery (placebo / random confounder / subset) | PARTIAL | DoWhy refuters genuinely coded in cell26 (run_refutation, PASS/FAIL/WARNING); no tests; not wired into cell36's shipped report path |
BUILD-small (B3.7) — unit tests around cell26's refutation assembly (mocked estimators); wiring into the shipped cell36 report path → BUILD_DEBT |
| A13 | Each conversion classified caused/anticipated, written to immutable ledger | PARTIAL→UNBACKED core | Lift machinery real (cell36 outcomes + CUPED reports, tested), but per-conversion caused/anticipated classifier and causal_credit_ledger exist nowhere (docs-only) |
BUILD (B3.2b) — rails classification job composing with M1 recompute (holdout-baseline → anticipated; incremental-above-baseline → caused), DDL unified.causal_credit_ledger (bitemporal, append-only), tests |
| A14 | Signal Factory desk runs full 7-stage SRPVDAL incl. deliberate guardrail block | BACKED (as demo) | mizoki_runtime/demo_signal.py 7 stages, seeded, guardrail veto; site tests |
NO ACTION — accurate as framed |
Copy consequence (owner decision D3): under the mission's build-to-claim policy, mechanism claims may stay present-tense once code exists and passes tests. Recommended one-line honesty addition to the §03 note (canon change, re-pin same commit): "Rails ship flag-gated and dry-run first; live posting is enabled per engagement." This keeps present-tense mechanism copy while making operational state honest (TRUTH.md 6.4). Live/scheduled operation remains operator-enabled per RUNBOOK — never claimed as running.
§2b — Dossier & site machinery claims (C19–C47)
Verdicts measured at 6ef0275, verified unchanged at e119a1f (A2 Task 3; full evidence in the A2 report). Resolutions reflect the owner's fix-all approval:
| ID | Claim | Verdict | Resolution |
|---|---|---|---|
| C19/C39 | Designs registered before exposure; "experiments on your spend" | PARTIAL (holdouts tested; ghost bids module-only; MDE absent) | B3 design_registry (mde + ghost_bid type); ghost-bid execution path → BUILD_DEBT |
| C20 | Reallocation reads the causal ledger | PARTIAL (ReLU-uplift scoring coded; ledger absent) | B3 builds unified.causal_credit_ledger; reallocation-reads-ledger wiring → BUILD_DEBT; B5 tests scoring constants |
| C21/C44 | Signal Factory / demo on live runtime | BACKED (demo engine, disclosed, tested) | No action |
| C22–C28, C31–C33, C35, C36 | Threshold binary-search agents; KG playbook writeback; $50 floor; 80/85/15% defaults; ReLU 5%/70%; ±20% clamps; five policies; z>2 fatigue; 90/10 holdback; diversity check; uplift quadrants | PARTIAL — all coded with constants matching page copy EXACTLY, none tested | B5 claims-backing test pack (new tests/claims_backing/, mocked externals) pins each coded constant/mechanic; ledger cites module+test per claim |
| C29 | CUPED-adjusted experiments | BACKED (cell36, tested) | Ledger cites |
| C30 | 10% always-on holdout | PARTIAL (core write-once 10% tested; budget-ladder enum-only) | Ledger cites cell36; B5 asserts ladder enum |
| C34 | Thompson-sampling bandits | BACKED (lift-engine + test) | Ledger cites |
| C37 | Qini/AUUC gates | BACKED (uplift_cohort_exporter min_qini + cell36 activation, tested) | Ledger cites |
| C38 | Customer Match / Meta CAPI audience export + writeback | PARTIAL (real sync clients, untested; no measured loop) | B5 payload-construction tests; e2e measured-writeback loop → BUILD_DEBT (activation stays uplift_export_cohort → guardrails → DCP) |
| C40 | Canonical event stream (Shopify→envelope) | BACKED (22 envelope tests + Shopify order tests) | Ledger cites |
| C41 | "Prove is live" | PARTIAL in-repo (machinery tested; liveness is a deploy fact) | Ledger cites live-verified memory records (cell36 live, IAM-locked 2026-07-31→08-07); no copy change |
| C42 | Intent API "served sub-100ms (Cell 28)" | PARTIAL (API is Cell 34, 60 tests; latency claim nowhere in code) | B2 fix T6 (design-target wording + Cell 34) |
| C43 | DoWhy stack "(Cells 26–27, 35)" | PARTIAL (X/DR+DoWhy real in cell26; shipped causal = 36) | B2 fix T7 → "(Cells 26–27, 35–36)" |
| C45 | "Native connectors 12 · Governed MCP tools 650+" | PARTIAL (9/12 in-repo code, 3 = MCP mounts; 650+ safe-true vs measured 1,137) | B2 adds GA4 tile → 13; ledger records basis |
| C46 | CMEK / VPC claims | CMEK UNBACKED (zero CMEK IaC); VPC PARTIAL (shared-VPC terraform real, no VPC-SC) | B2 addendum: reword CMEK to roadmap/design-target framing, VPC wording to shared-VPC truth; CMEK IaC → BUILD_DEBT (operator/terraform — out of agent scope) |
| C47 | Tamper-evident audit ledger | BACKED (hash-chained store + verify_chain, tested, served by service-audit-replay) | Ledger cites |
Systemic finding: every "operating default" number on the dossiers matches its coded constant exactly — the gap was tests, closed by B5. Copy gaps concentrate in C42/C43/C46, closed by B2.
§3 — Content & truth-discipline fix list (Phase B2 scope)
Served-page findings from A1 (full text in A1 report §2). All these pages except blog/* are canon-pinned → edits require this checkpoint's owner approval and a same-commit canon.lock.json re-pin (scripts/check_design_canon.py --update).
| # | File:line | Finding | Proposed fix |
|---|---|---|---|
| T1 | blog/decision-control-plane.html:519 |
"89% of decisions autonomously / 11%" unlabeled, presented as capability fact | Label as design target ("designed to route ~89%… — design target") |
| T2 | index.html §02 #pipeline (532–644) |
ACT-991 sim figures ($5.0M/DEL 39/94%/98%/12%) lack an in-section ILLUSTRATIVE cover (only hero-strip + #control have one) | Add in-section "ILLUSTRATIVE SCENARIO" marker to #pipeline (TRUTH 2.5) |
| T3 | index.html:488 |
Banner: "REPRESENTATIVE OF A DEPLOYED MIZOKI3 SYSTEM" vs ceiling "built, pre-benchmark" | Reword: "REPRESENTATIVE OF A FULLY CONFIGURED MIZOKI3 SYSTEM · ILLUSTRATIVE FIGURES" |
| T4 | executive-briefing/js/data.js:331,391,452 |
"pilot units / pilot lines / pilot book" metric rows with no pilot readout (TRUTH 2.3) | Reword rows to composite framing ("in composite scenario units") or add "(composite)" per row |
| T5 | data.js per-row metrics (82–514) |
Figures gate-pass via section label but lack per-figure labels | Optional: add "(composite)" to each row — low priority, include if cheap |
| T6 | demo-signal.html:405 |
"served sub-100ms (Cell 28)" — design target stated as operating behavior + stale cell id (shipped scoring = Cell 34); invisible to gate regex | "designed for sub-100 ms response (Cell 34) — design target"; extend gate regex to sub-\d+ms |
| T7 | demo-signal.html:406 |
"(Cells 26–27, 35)" plan-vintage causal stack numbering (shipped causal = 36) | Correct to shipped numbering (26–27, 35–36) |
| T8 | demo-capital.html:251 (+meta :10) |
15% headroom threshold unlabeled (same number labeled illustrative on demo.html) | Add "illustrative scenario threshold" wording to match hub |
| L1 | security.html:93 (dead-routed) |
"Guaranteed Rollbacks" heading | Reword "Verified Rollbacks" (rollback drills are the real mechanism) |
| L2 | roi.html:769 (dead-routed) |
"50-75× improvement" unlabeled | Label design target or remove |
| K1 | pricing.html:165,167 |
/#platform, /#governance anchors don't exist on index |
Point at real ids (#what / #control) or add ids |
| K2 | blog/posts.json |
Two og images missing on disk (feed emits 404 URLs) | Generate the two og.png assets (or point at existing art) |
| K3 | executive-briefing/js/app.js:897 |
Fallback #contact anchor doesn't exist (latent, embed-only) |
Change fallback to /contact |
| W1 | index vs signal/pricing | L0 and L1 both marked "Default."; pricing "L0–L1 · approval-first" collides with L2 rung name "Approval flow" | Make L0–L1 "Starting rungs · Default posture" (match signal.html); pricing → "L0–L1 · observe & recommend" |
| W2 | Connector lists | GA4 first-class in demo + measurement rails but absent from homepage 12-connector grid; Meta descriptor drift | Owner call D4: add GA4 tile (grid becomes 13 — stat row updates) or leave; align Meta descriptor |
A3 cross-reference: story bank v1.1 and net-yield positioning identical everywhere (no action). Live == repo on all audited routes; banned-string scan of live pages clean (only negations).
§4 — Discipline machinery (Phase B1 scope, revised)
content_qa.py exists, self-tests, and gates deploy-homepage. Gaps to close:
- Unbacked-claims check (new rule E): a machinery-claims ledger (
# MIZ OKI 3.5/docs/marketing/claims-ledger.yaml) mapping each scoped machinery claim (M1–M9, A10–A13, C-rows that stay present-tense) → eitherbacked:(list of code+test paths that must exist) ordebt:(BUILD_DEBT.md entry id). content_qa fails if a ledgerbackedpath is missing on disk, if adebtid is absent fromdocs/BUILD_DEBT.md, or if a scoped page carries a machinery-claim marker not in the ledger. Self-test seeds one violation per sub-class. - Scope expansion: add all blog posts, demo-capital/estate/risk/nexus, pricing.html, security.html, roi.html, index.html (B2 fixes those pages first in the same branch so the expanded gate lands green; false-positive tuning for CSS/JS numerics as A1's T9 showed).
- Regex hardening:
sub-\d+msjoins the sub-second/15-minute class. docs/BUILD_DEBT.mdcreated with the approved debt rows (each: claim, page anchor, owner note, flip condition).- skills_sync: no work (passes).
- No
.github/workflows/**edits in this run (HTTPS push lacks workflow scope; content_qa is already wired into the deploy gate, which is the gate that matters).
§5 — Merge-safety map (measured, A2 §g)
- Auto-merge bot:
claude|gemini|codex|copilot|cursor/**push → merges tomainin seconds (-X theirsretry; protected-path exception opens a review PR for^\.github/|^deployment/terraform/|^deployment/cloudbuild|^CODEOWNERS$). - Deploy Router fires only
deploy-*.ymlwhoseon.push.pathsmatch the merge diff. Measured:services/net-yield/**fires nothing (deliberately router-inert); a newservices/measurement-rails/**fires nothing;# MIZ OKI 3.5/**fires nothing (site is dispatch-only). B3 landing paths are chosen to keep the whole diff router-inert; builders must re-verify each final path against everydeploy-*.ymlbefore commit. ci.yamlis PR-only (lint, consolidated tests, skills_sync, selfcheck) — opening the PR gives free validation; the auto-merge lane runs none of it (another reason to avoid the auto-merge lane).- Therefore: branch
completion-run(never auto-merged) + draft PR + owner-clicked merge is the only flow that satisfies "nothing lands on main before my approval". Post-merge, this diff auto-deploys nothing; site changes go live only via a later owner-dispatcheddeploy-homepage.ymlwithAPPROVED.
§6 — Proposed Phase B build list (for Checkpoint A approval)
- B3 — Measurement Rails (the core): new
services/measurement-rails/(router-inert; net-yield service pattern: transport-injected clients, flag module, pytest ≥80% on business logic, mocked platform APIs only): B3.1 house recompute +unified.house_attributionDDL + lag profiles; B3.2b caused/anticipated classifier +unified.causal_credit_ledgerDDL; B3.3a–e five rails (Google EC, Meta CAPI 48h, AEM config, GA4 MP validate-first, offline conversions 90-day) — all flag-gated, DRY_RUN default (build/validate/log, never send); B3.4 drift monitor (3-day logic, Prometheus metric + Grafana rule stub); B3.5 writeback module behind hardMEASUREMENT_WRITEBACK=falsepinned by a test; B3.6 MDE field + ghost_bid design type in the experiment registry; B3.7 cell26 refutation-assembly tests;docs/measurement-rails/{RUNBOOK,ADR-MR-001}.md. Boss legacy modules are NOT edited (their deploy paths would auto-redeploy the boss on merge); they are recorded in BUILD_DEBT as migration/retirement items. - B2 — Content hardening: §3 fix list (T/L/K/W rows) + §03 honesty line (D3) + canon re-pin same commit; per-dossier story excerpts produced as an approval-gated diff only (
docs/completion-run/dossier-stories.proposal.diff— not applied): Doorman→audiences, CFO/brand-search→budget, Coupon→budget, Prove-It-Or-Lose-It→measurement, Bundle→measurement. Briefing restructure: already applied 2026-08-07 — no new proposal. - B1 — Discipline machinery: §4 items (claims ledger + rule E + scope expansion + BUILD_DEBT.md), sequenced after B3 so ledger
backed:paths cite landed code. - B4 — Net-yield: SKIPPED — complete (A2 §d). Preview labels stay per claim ledger.
§7 — Decisions requested at Checkpoint A
- D1: Approve the build list above (including the BUILD vs BUILD_DEBT splits in §2/§2b).
- D2: Branch =
completion-runpushed with a draft PR (recommended; the session-assignedclaude/…name would auto-merge to main on push, violating mission rule 2). - D3: §03 note honesty line on /signal/measurement (canon edit) — yes/no.
- D4: GA4 homepage connector-grid tile (grid 12→13) — add or leave.