A2 — Platform & Machinery Audit (MIZOKICloudRun)

Auditor: Subagent A2 (read-only). Repo: /home/user/MIZOKICloudRun, branch claude/mizoki-completion-run-4ms74o == origin/main @ 6ef0275. Date: 2026-08-08. Method: every verdict below is from direct file reads / grep / read-only command runs at this commit. Nothing was edited. Test counts are pytest --collect-only -q where the sandbox could collect, plus grep -c "def test_" for files whose collection failed on a sandbox-missing fastapi module (noted inline — the errors are environment-dependency, not code, failures). Verdict rubric (as tasked): BACKED = working code + tests exist for the mechanism as claimed · PARTIAL = some code exists but incomplete/stub/different scope · UNBACKED = no code.

Headline correction of the prior brief: # MIZ OKI 3.5/scripts/content_qa.py, the net-yield stack, and scripts/skills_sync.py all EXIST and pass their own checks at this commit (evidence in §a, §b, §d). The prior task brief claiming they are missing is stale.


TASK 1 — Machinery-claim verdicts (/signal/measurement)

Claims source: # MIZ OKI 3.5/signal-measurement.html (311 lines). M1–M4 from §01 (lines 226–236), M5–M9 from the §02 rails table (lines 244–251), A10–A13 from §03 (lines 260–265), A14 from §04 (line 277). Page carries its own honesty note at line 267 ("mechanisms on this page are operating machinery, and the windows and thresholds shown are operating defaults").

# Claim Verdict Evidence
M1 House attribution recompute (7-day click / 1-day view, nightly job, BigQuery unified.house_attribution, per-source lag profiles) PARTIAL Module exists: miz-oki-adk-agents/boss/cross_platform_attribution_integration.py — run_attribution_recompute() (:184-293) does a Pandas 7-day-window time-decay recompute and writes a Firestore weights collection; MCP tool attribution_recompute registered (:478-485) and mounted at boss startup (boss_agent_core.py:39555-39558, ENABLE_CROSS_PLATFORM_ATTRIBUTION default true :1845). But: the data fetch is a hard-coded mock — _fetch_recent_data (:402-428) generates np.random dummy frames ("In production, retrieves from BigQuery/Firestore"); there is NO 1-day-view logic (only one 7-day filter, :227-230); roas_platform/roas_model are written as 0.0 placeholders (:262-263); no nightly scheduler anywhere (no hits in jobs/, data-pipelines/, deployment/, gitops/); house_attribution appears nowhere in the repo (zero grep hits — no BQ table/DDL); lag profiles exist only as a static, never-loaded config config/cross_platform_attribution/lag_profiles_default.json ("sampleSize": 0, "initial estimates"; zero code references to the filename) — the recompute uses fixed 1/(hours+1) decay, not per-source lag profiles. Tests: only the tool-name registry contract in tests/test_integration_srdal_pipeline.py:314-316; no functional tests.
M2 Salted SHA-256 identity stitching, deterministic-only for causal math, raw identifiers never persisted BACKED (with two labeled nuances) Real cell + tests: src/cells/identity_attribution/identity_cell/crypto/hashing.py — peppered SHA-256 (sha256_token :28-31 mixes a KMS-managed pepper; normalization for email/phone/etc.), rule ":no raw PII may leave the connector" (:3-6); privacy gate scans and rejects raw emails/phones (pipeline/privacy_gate.py); deterministic+probabilistic resolver (resolver.py). Tests: tests/test_crypto.py (10), test_privacy.py (5), test_resolution.py (10). Auto-deploy wired: .github/workflows/deploy-identity-attribution.yml (push paths src/cells/identity_attribution/**). Deterministic-only causal gate: src/cells/cell36/causal_cell/holdouts.py:182-194 REFUSES identity_kind != "deterministic" ("probabilistic identities never [enter causal math]"), salted deterministic-hash arm assignment (assignment_bucket :59-64, INTENT_HOLDOUT_SALT in config.py:53); tests src/cells/cell36/tests/test_holdouts.py (8 collected). Nuances to keep labeled: (1) technically a global KMS pepper, not a per-record salt — identity_cell/config.py:46 says so explicitly; (2) a parallel, weaker stitcher exists in cross_platform_attribution_integration.py:_hash_identifier (:118-124) that is plain UNSALTED sha256(email) into Firestore.
M3 Drift monitor (platform vs model revenue, >20% for 3 consecutive days, Prometheus/Grafana) PARTIAL (stub) Only code: cross_platform_attribution_integration.py — DriftAlert dataclass (:74-83), check_drift(threshold_pct=20.0) (:373-391) whose body is a stub returning {"alerts": [], "message": "Drift check logic requires spend data synchronization (future enhancement)"}; MCP tool attribution_drift_check registered (:497-504). No 3-consecutive-days logic anywhere (repo grep for revenue-divergence consecutive logic: none — slo_monitor.py/circuit breakers are unrelated). No Prometheus/Grafana wiring for it: monitoring/alerts.yaml is Cloud Monitoring error-rate policies for cells (no drift/attribution/revenue entries); monitoring/grafana/datasources/ holds only a prometheus datasource. Only roas_platform/roas_model fields in the whole codebase live in this one module (mock-fed).
M4 Weight back-propagation of corrected conversion values to Meta/Google bid systems (flag-gated) PARTIAL Two half-mechanisms: (1) boss push_weights_to_platforms (cross_platform_attribution_integration.py:299-367) reads Firestore weights and generates increase/decrease actions, but _execute_platform_adjustment (:430-433) is a log-only placeholder ("Would use UnifiedPlatformIntegration clients here") and it is NOT flag-gated (only a dry_run param). (2) The real flag-gated corrected-value writeback is net-yield scope: services/net-yield/writeback/meta_capi.py (net-contribution CAPI Purchase values, SHA-256 external_id, refuses incomplete economics :39-43) and google_value_rules.py (RESTATEMENT conversion adjustments + recommend-only ConversionValueRule multipliers), both hard-gated by NET_YIELD_WRITEBACK default false (writeback/__init__.py:24) with send_* raising WritebackDisabled; tests pin the gate (test_writeback.py::test_flag_defaults_off, test_meta_send_refuses_while_off, test_google_send_refuses_while_off, + payload tests; 12 tests). Both modules are payload construction only — transports injected, no scheduler/actuator, self-labeled "claim_label: built, pre-benchmark", L1 recommend-only. End-to-end "corrected attribution flows to bidders" does not run anywhere.
M5 Google Enhanced Conversions server-side rail PARTIAL Substantial client code, no tests: miz-oki-adk-agents/boss/enhanced_conversions_integration.py (1,712 lines) — PIIHasher (:236-378; SHA-256 with Google-spec normalization of em/ph/fn/ln/ct/st/zp/country), GoogleEnhancedConversionsClient (:381-580; OAuth refresh-token flow, uploadConversionAdjustments on googleads.googleapis.com/v22 :77, batch 2000, retries w/ backoff), FastAPI endpoints + MCP tools registered in boss core (boss_agent_core.py:797-812, registration :26493; ENABLE_ENHANCED_CONVERSIONS default true :1803). No dedicated unit tests anywhere (only the tool-name registry list in tests/test_integration_srdal_pipeline.py); no scheduled pipeline feeding it; credentials all env-var defaults "". Related-but-different: net-yield google_value_rules.py (adjustment payloads, gated — see M4).
M6 Meta CAPI, shared event_id + 48h dedup PARTIAL MetaConversionsAPIClient in enhanced_conversions_integration.py:586-895 — real httpx client to graph.facebook.com/v25.0/{pixel_id}/events, event_id carried on every payload (:745), Firestore + in-memory event-id dedup (_check_duplicate :1015-1041). The coded dedup window default is 24h, not the claimed 48h ("window_hours": 24 :106; no 48-hour dedup constant anywhere in the repo). Second uploader with quota governance: conversion_tracking_integration.py:1825-1908 (MetaCAPIUploader, X-App-Usage header handling) — its payloads carry no event_id field. Also: services/net-yield/writeback/meta_capi.py (deterministic event_id=f"netyield-{tenant}-{order}" :47, gated); services/ekis/src/connectors/facebook/fb.graph.client.ts (CAPI /events); src/cells/cell02/integrations/clients/facebook.client.ts (sends event_id :133). No tests for any CAPI path (ekis tests cover only klaviyo/redis/ratelimit). Browser-pixel counterpart that would share the event id is not in this repo.
M7 AEM 8-event priority schema config PARTIAL Coded in the Meta client, not externalized config, no tests: enhanced_conversions_integration.py — AEM_EVENT_PRIORITIES (11 events, priority 1–8, :595-607), AEM_MAX_EVENTS = 8 (:610), configured-events list sliced to 8 (:617-620), is_aem_event/get_aem_priority/filter_aem_events (:622-713), LDU data-processing options (:655-680), options attached on send (:777-785). It is priority filtering/annotation in one module; there is no per-tenant AEM schema config file and no test exercising it.
M8 GA4 Measurement Protocol, validate-before-send PARTIAL Three MP clients exist: (1) conversion_tracking_integration.py:488-620 GA4MeasurementProtocolClient — COLLECT_URL + DEBUG_URL (/debug/mp/collect :502), but debug is a constructor-time mode switch (debug XOR collect), not validate-then-send; (2) platform_compliance_integration.py:133-368 GA4MeasurementProtocolEU — per-call debug flag + EU regional endpoints (:59-63); (3) services/ekis/src/connectors/ga4/ga4.measurement.ts — separate validateEvent() (debug endpoint, :110-139) and sendEvent() (:82-105); sendEvent does NOT call validateEvent first (validation used in the health-check path :338). Also GA4 debug-endpoint credential ping in services/service-marketing-connectors/connector_credentials.py:454-460. So: validation capability everywhere, validate-before-send pipeline nowhere; no tests on any GA4 MP client.
M9 Offline conversions upload (GCLID/WBRAID/GBRAID, 90-day window) PARTIAL Three uploaders: (1) most complete — services/ekis/src/connectors/googleAds/ads.conversions.ts (GoogleAdsConversionsClient on the google-ads-api npm lib; OfflineConversion supports gclid/gbraid/wbraid :19-27, operation builder prefers gclid→gbraid→wbraid :230-235, real customer.conversionUploads.uploadClickConversions :106, partial_failure); no tests for it (ekis test dir covers other connectors). (2) src/cells/cell02/integrations/clients/googleads.client.ts:93 (uploadClickConversions) + tested ID extraction src/cells/cell02/tracking_id_extractor.py (gclid/wbraid/gbraid; tests/test_tracking_id_extractor.py, 3 tests). (3) boss conversion_tracking_integration.py:1916-2030 GoogleAdsOfflineUploader — the Google Ads client import is commented out (:1929-1934: # from google.ads.googleads.client import GoogleAdsClient … self.service_available = False), so this path ALWAYS runs "Mock Mode" (:2005-2013); GCLID only. No 90-day-window enforcement anywhere (zero hits in all three). MCP tool google_ads_upload_conversion registered (:2172).

Additional machinery claims found on the page (§03 incrementality engine, §04 factory)

# Claim (page text) Verdict Evidence
A10 "Design registered before the first impression — randomized holdout, ghost bids…, or matched control geographies. Minimum detectable effect declared up front." (:261) PARTIAL Holdouts are real + tested: cell36 write-once deterministic-hash holdouts (holdouts.py; POST /v1/causal/holdouts:assign; 10% permanent, INTENT_HOLDOUT_FRACTION/INTENT_HOLDOUT_SALT in config.py:48-58; tests test_holdouts.py + test_causal_api.py). Experiment registry with geo/holdout/rct designs: services/service-media-incrementality/main.py:40-52 (148-line registry service, measurement hierarchy :23-29). Ghost bids: dataclasses/decision fields only, in miz-oki-adk-agents/boss/autonomous_budget_reallocation_mvp.py:163-205 (+ eshkg_signal_integration.py, boss_agent_core.py) — no ghost-bid auction logger service. MDE nowhere in code (zero grep hits for minimum_detectable/mde in src/ + services/) — matches the active-memory record that the plan-vintage /api/v1/causal/holdout {…mde} shape "does not exist on the shipped surface".
A11 "Heterogeneous effects estimated per segment and region with meta-learner models" (:262) PARTIAL Real meta-learner code: src/cells/cell26/Cell26.py — causalml XLearner/DRLearner/S-Learner CATE per edge (:24-25, :124-128, :196-283); cell36 estimators.py reuses DRLearnerWeightTrainer + CUPED with recorded provenance/limitations (:63-84) and tests (test_causal_api.py, 19 test fns). But cell26 itself has no tests (dir = Cell26.py + Dockerfile + requirements only) and no deploy workflow references it besides fleet scripts.
A12 "Automated refutation… placebo treatments must collapse the effect, random confounders must not move it, subset re-runs must agree" (:263) PARTIAL DoWhy refutation is genuinely coded in src/cells/cell26/Cell26.py: run_refutation (:293-…) with placebo-treatment, random-common-cause, bootstrap refuters and refutation_overall: PASS/FAIL/WARNING (:176-177), run_refutation: bool = True (:134). No tests for it; not wired into cell36's shipped report path (cell36 uses CUPED/DR without the DoWhy refuter battery).
A13 "Each conversion is classified caused or anticipated… written to the immutable ledger" (:264) PARTIAL Lift machinery exists (cell36 /v1/causal/outcomes bitemporal + /v1/causal/reports:run default CUPED, tested), but per-conversion caused/anticipated classification exists nowhere in code: unified.causal_credit_ledger and lii_get_causal_credit appear only in skill docs (src/shared/virtuoso_models/skills_data/SKILL.md:1490, boss_agent_skillpack.md:33); zero code/SQL hits for causal_credit_ledger or a caused/anticipated classifier. The on-site "ledger" exhibit is illustrative (page's own note, line 267).
A14 "The Signal Factory desk runs the full seven-stage SRPVDAL loop on live runtime — including a deliberate guardrail block" (:277) BACKED (as the demo engine it is) # MIZ OKI 3.5/mizoki_runtime/demo_signal.py — 7-stage STAGES = ("sense","reason","plan","validate","decide","act","learn") (:489), deterministic seeded runs, deliberate guardrail veto; served by real Flask routes; site tests tests/test_demo_signal.py + test_demo_platform.py. Accurate as stated for the demo desk (deterministic fixture on the site's own runtime, not platform cells) — keep that framing.

Cross-cutting Task-1 summary: the marketing page's §01–§02 mechanisms trace to one untested boss integration module family (cross_platform_attribution_integration.py, enhanced_conversions_integration.py, conversion_tracking_integration.py, platform_compliance_integration.py) plus scattered TS connectors (ekis, cell02) — mostly payload-building clients, none exercised by tests, several with mock/stubbed cores; the claimed BQ table, nightly job, drift alerting, 48h window, 90-day window, and salish "salted" cross-platform stitcher do not exist as claimed. §03's engine is the strongest area (cell36 + cell26 + service-media-incrementality), but MDE registration, DoWhy-refutation-in-the-shipped-path, ghost-bid logging, and the caused/anticipated per-conversion ledger are not shipped. "Code that builds payloads for X" ≠ "the claimed pipeline end-to-end" applies to M4–M9 across the board.


TASK 2 — Platform state report

a) # MIZ OKI 3.5/scripts/content_qa.py — EXISTS, runs clean

b) skills_sync + parity — PASSES

c) LII cells + MCP

d) Net-yield — COMPLETE per the checklist; nothing missing

e) docs/marketing (in # MIZ OKI 3.5/docs/marketing/)

f) Doorman / briefing acts / demo-hub lines

g) CI configuration (merge-safety for the build phase) — PRECISE FACTS

h) Reusable Google Ads / Meta connector assets for measurement rails

i) MEASUREMENT_WRITEBACK flag — DOES NOT EXIST (verified)

Repo-wide grep for MEASUREMENT_WRITEBACK: zero hits. The only writeback flag today is NET_YIELD_WRITEBACK (services/net-yield/writeback/init.py, default false, test-pinned).

j) docs/BUILD_DEBT.md and docs/measurement-rails/ — DO NOT EXIST (verified)

ls both paths: No such file or directory; repo-wide find for *BUILD_DEBT*: zero hits. docs/measurement-rails/: absent.


Notes for the report consumer (truth-discipline compliant)


TASK 3 — Dossier & site claims C19–C47

Same rubric as Task 1 (BACKED = code + tests for the mechanism; PARTIAL = code exists but incomplete/stub/untested/different scope; UNBACKED = no code), plus DEMO-ONLY where the only backing is the site's own deterministic demo runtime (# MIZ OKI 3.5/mizoki_runtime/demo_*.py) — per SKILLS_V3_ROLLOUT_REPORT.md note 5, those engines are "deterministic stdlib fixtures (no LLM calls)". All findings measured at commit 6ef0275; "registered/coded" never means live-verified.

signal.html

# Claim Verdict Evidence
C19 "Signal runs the experiment — holdouts, ghost bids, matched geographies" (signal.html:177; og :10 "Signal runs the experiment") PARTIAL Holdouts: BACKED slice — cell36 write-once deterministic holdouts + outcomes + reports (src/cells/cell36/causal_cell/holdouts.py, tests test_holdouts.py/test_causal_api.py). Ghost bids: coded in-module only — miz-oki-adk-agents/boss/autonomous_budget_reallocation_mvp.py GhostBidTest engine (create/assign/results :953-1245, ghost_validated = p_value < 0.05 :1940, 2% default holdout :254); no service, no tests. Geo: registry-only — services/service-media-incrementality/main.py:40-52 (design ^(geo|holdout|rct)$), no estimator pipeline. No end-to-end "experiments on live spend" pipeline exists; MDE registration absent repo-wide (zero minimum_detectable/mde hits in src/+services).
C20 "Budget reallocation reads from this ledger — never from platform-reported ROAS alone" (signal.html:241) PARTIAL The functional rule is coded: reallocator scores ReLU(delta) × confidence × log(sample_size) from uplift edges (autonomous_budget_reallocation_mvp.py:233, ReLUGatingOptimizer :1293-1345) and prefers holdout/ghost-validated lift (:1349, :1940); iROAS-not-platform-ROAS is platform law (CLAUDE.md §3). But the named artifact — the immutable causal-credit ledger — does not exist (Task 1 A13: zero causal_credit_ledger code/SQL hits); the module reads Firestore/KG edges and is untested.
C21 "The Signal Factory desk runs the full loop live" (signal.html:281; §04 of signal-measurement) BACKED — DEMO-ONLY # MIZ OKI 3.5/mizoki_runtime/demo_signal.py — 7-stage STAGES tuple (:489), canonical-event conversion (:315), deliberate guardrail veto; site tests tests/test_demo_signal.py + test_demo_platform.py. True of the site's own deterministic, seeded runtime; not platform cells.

signal-thresholds.html

# Claim Verdict Evidence
C22 Binary-search threshold discovery agents (":256-263 — discovery agent runs a binary search…never probe below 50%") PARTIAL miz-oki-adk-agents/boss/relu_threshold_agents.py — ThresholdDiscoveryAgent binary search (:213-398; start 150% :62, floor 50% min_bid_multiplier: 0.5 :63, tolerance 5% :64, max 10 iterations :66); registered in boss core (:1428-1435). No tests (only the tool-name registry list). Older parallels: srpaldl-cells/planner/app.py:59, scripts/py_scripts/main.py:373.
C23 "Write the discovered threshold to the knowledge graph so every later campaign in the vertical starts warm" (:264, :316) PARTIAL Persistence exists but is Firestore, not the Neo4j KG: per-industry pattern library relu_playbook_patterns (relu_threshold_agents.py:872-950), searches stored in relu_threshold_searches (:427). Untested; "knowledge graph" is generous for a Firestore collection store.
C24 "$50/day floor per ad set" (:277-278, labeled default) PARTIAL Exact coded default: "min_budget_per_adset": 50.0 (relu_threshold_agents.py:70; module doc :17 "Minimum $50/ad set"). Label-accurate, but no test pins it.
C25 "80% of budget flows to the top performers (operating default)" (:282) PARTIAL "concentration_factor": 0.8 # 80% budget to top performers (relu_threshold_agents.py:71; SignalAmplificationAgent :442-488). Label-accurate coded default; untested.
C26 "85% exploit / 15% explore (operating defaults)" (:296) PARTIAL "exploitation_rate": 0.85 # 85% exploitation, 15% exploration (relu_threshold_agents.py:77). Label-accurate coded default; untested.

signal-budget.html

# Claim Verdict Evidence
C27 ReLU gate — score = ReLU(uplift)×confidence×log(n); uplift ≥5%, confidence ≥70% (:225-229) PARTIAL Exact formula + exact gates coded: autonomous_budget_reallocation_mvp.py DEFAULT_THRESH_UPLIFT=0.05 / DEFAULT_THRESH_CONF=0.70 (:102-103), gated_score = relu_delta * confidence * math.log1p(sample_size) (:233), ReLUGatingOptimizer.gate_edges (:1293-1345); twin module ad_budget_reallocation_mvp.py:386; deployed sibling services/relu-evaluation-service/src/relu/{evaluator,policies}.py (uplift scoring service; no tests found); demo mirror mizoki_runtime/demo_signal.py:384. No tests on any of them.
C28 Pacing clamps ±20%/period; daily cap 10%; >20% routes to a human (:242-243, :284-285) PARTIAL DEFAULT_DAILY_CAP=0.10 / DEFAULT_MAX_STEP=0.20 (:104-105); pacing rule spend_t = spend_{t-1} × clamp(1 + β(uplift/target−1), 0.8, 1.2) in ReLUGatingOptimizer (:1298-1302); ">20% requires approval" (:2575); hard 10% cap (:1535). Untested module.
C29 "CUPED variance reduction" feeds the uplift (:251) BACKED uplift_pacing_integration.py _cuped_estimate (:446-455; UpliftMethod.CUPED :58, default :300) reused by cell36 estimators.py (:74-84) — and TESTED: src/cells/cell36/tests/test_causal_api.py:108 asserts method=="cuped", :113-116 asserts provenance incl. the known limitation variance_reduction=0.4 (fixed theta, not computed from pre-period), :365 cuped_available. Keep the fixed-theta limitation labeled.
C30 "Canary 10% → expansion 20% → full 70%; permanent 10% holdout the engine can never touch" (:294-295) PARTIAL Tested 10% permanent holdout exists on the intent-causal side: cell36 DEFAULT_HOLDOUT_FRACTION = 0.10 (config.py:28, env :49) with write-once semantics (test_holdouts.py). The budget engine's own ladder is enum-level only: RolloutStage CANARY/…/HOLDOUT "10% permanent holdout" (autonomous_budget_reallocation_mvp.py:147-150), ghost default 2% (:254); untested, no enforcement pipeline.
C31 Five always-on policies with cooldowns — CPA Guard, ROAS Accelerator, Fatigue Swap, New-Segment Probe, Margin-Aware Bid (:305-315) PARTIAL Exactly these five are declared in miz-oki-adk-agents/boss/policy_engine_integration.py — p1_cpa_guard :1034, p2_roas_accelerator :1067, p3_fatigue_swap :1099, p4_segment_probe :1131, p5_margin_bid :1163 — on a declarative condition→action→cooldown engine (:275-289; 1,907 lines). No tests.

signal-creative.html

# Claim Verdict Evidence
C32 Fatigue triggers: z-score > 2 vs 7-day baseline, confirmed on two consecutive windows (:224-227) PARTIAL miz-oki-adk-agents/boss/creative_fatigue_integration.py — z-score thresholds 2.0 for CTR/CVR/CPA (ctr_zscore_threshold etc. :278-280), FatigueSignalCheck.z_score (:169-181), status ladder (:75-79). Untested (registry list only).
C33 Holdback rotation 90/10 (:250-253 "90% moves to the new creative; 10% stays…operating default") PARTIAL RotationStrategy.HOLDBACK = "holdback" # Keep 10% on old for comparison (creative_fatigue_integration.py:98; RotationExecutor per module doc :13). Untested.
C34 Thompson-sampled selection; bandit creative tests, UCB where regret matters (:255-257, :270) BACKED Tested implementation in lift-engine: services/lift-engine/src/core/dynamic_uplift_rl.py ThompsonSamplingBandit (:187-203) + UCB strategy (:281-312), exercised by services/lift-engine/tests/test_causal_reasoning.py:441 (strategy='thompson_sampling'). Creative-side Thompson also coded (untested): creative_fatigue_integration.py:585-651 _thompson_sample; legacy src/cells/cell27/Cell27.py:77,:509 (no tests, .final* file clutter).
C35 Semantic-diversity check refuses near-duplicates (:259) PARTIAL creative_fatigue_integration.py — embedding field (:224), min_semantic_distance threshold (:302, :592), diversity filter before selection (:611). Untested.

signal-audiences.html

# Claim Verdict Evidence
C36 Uplift quadrant segmentation — persuadables / sure things / lost causes / sleeping dogs (:221-241) PARTIAL miz-oki-adk-agents/boss/persuadability_segmentation.py — PersuadabilitySegment enum with the four quadrants (:48-53), rule-ordered classification thresholds (:72-80), per-segment treatment actions (:101). Also journey_uplift_counterfactual.py, causal_decision_systems.py. No tests.
C37 Statistical activation guardrails — "Qini coefficient ≥ 0.02: the model must beat random targeting" (:254) BACKED Exact gate coded: miz-oki-adk-agents/boss/uplift_cohort_exporter.py:147 "uplift_quality": {"min_qini_score": 0.02, "min_auuc": 0.55, "min_ate": 0.005} enforced at :178-181; shipped activation path carries qini/auuc: src/cells/cell36/causal_cell/activation.py:103-104 → uplift_export_cohort → guardrails → DecisionControlPlane (:25); TESTED — test_causal_api.py:110-111 asserts qini_score/auuc in the report, :312-317 asserts they flow into the activation prepare payload.
C38 Cohorts export to Google Customer Match / Meta Custom Audiences, predicted iROAS before, measured written back after (:260) PARTIAL Real REST clients, untested, and no closed measurement-writeback loop: unified_platform_integration.py sync_customer_match (offlineUserDataJobs create→addOperations→run, :725-815), sync_custom_audience (:1149), dual-platform sync (:1563-1580); export guardrails per C37. "Measured version written back after" exists nowhere as a pipeline.

shopify.html

# Claim Verdict Evidence
C39 "§04 What you get today … Holdouts, ghost bids, and matched-market experiments on your actual spend" (shopify.html:170-177) PARTIAL Same machinery as C19 (cell36 holdouts tested; ghost bids module-only; geo registry-only). The "today / your actual spend" present-tense framing rests on dispatch-only deploys (deploy-intent-platform.yml) and shadow-mode records — nothing in-repo demonstrates experiments running on customer spend. Copy-risk flag for the owner rather than a code gap alone.
C40 "One canonical event stream — Shopify orders, Google Ads, Meta, and lifecycle email emit the same governed business event" (:180-182) BACKED (mechanism; per-source depth varies) The envelope is real and tested: contracts/canonical-event-envelope/ (schema + 22 tests) and contracts/mizoki_contracts/envelope.py (CanonicalEventEnvelope). Shopify orders→envelope: services/net-yield/order_economics.py:294-305 build_envelope ("no second envelope"), asserted in test_order_economics.py:116; live Shopify intake services/intent-shopify-extender (HMAC, 13 tests). Google Ads/Meta→journey envelope: services/gemini-kg-pipeline/src/journey_events/{envelope.py,meta_connector.py,meta_worker.py} (own deploy workflow); pull adapters emit CanonicalRecords (service-marketing-connectors/direct_connectors.py); services/service-canonical-ingestion/ exists. Lifecycle email (Klaviyo) is the thinnest leg (ekis klaviyo.ts + credential spec).
C41 "Prove is live in the ledger above" (roadmap §, :197-198; Profit/Anticipate carry Preview tags) PARTIAL PROVE machinery is the platform's best-backed slice (cell36 holdouts/outcomes/reports, tested — Task 1 M2/A10/C29/C37), and repo records document measured shadow-deploy service URLs (SKILLS_V3_ROLLOUT_REPORT.md:160-165, README Live proofs §). But "is live" is a deployment/live-state claim this audit cannot verify from the repo, the intent-platform deploy is dispatch-only, and "the ledger above" (:135-141) is a labeled composite scenario, not customer output. The immediate neighbors keep Preview framing (:197, :215) — this one sentence carries the strongest unhedged present tense on the page.

Demo pages / index.html

# Claim Verdict Evidence
C42 "Intent Scoring API — …served sub-100ms (Cell 28)" (demo-signal.html:405, inside the ORACLE section that carries "Preview · in development" :397) PARTIAL — two copy errors The scoring API exists and is tested, but it is cell34, not Cell 28: src/cells/cell34/scoring_cell/ (main/scoring/orchestrator; 60 tests) — SKILLS_V3_ROLLOUT_REPORT.md:166: "cell 28 remains a legacy Cell28.py cell; scoring lives in cell 34"; src/cells/cell28/Cell28.py is the Sales Revenue Pipeline. And "served sub-100ms" appears nowhere in cell34 code, SLOs, or measurements (zero latency/p95 hits in scoring_cell) — a design target phrased as observed serving, which the truth discipline bans. Copy should read cell 34 and label the latency a design target.
C43 "X-Learner / DR-Learner uplift and DoWhy refutation… (Cells 26–27, 35)" (demo-signal.html:406) PARTIAL — wrong cell map The mechanisms exist: cell26 has causalml X/DR/S-learners + DoWhy placebo/random-common-cause/bootstrap refuters (Cell26.py:24-25, :124-128, :293-…, untested), and the shipped causal surface is cell36 (DR/CUPED, tested). But the page's mapping is inaccurate: cell27 is Thompson-sampling online decisions (legacy, untested), cell35 is the LII intent graph cell (not causal refutation), and cell36 — the actual shipped causal cell — is unmentioned.
C44 Demos run on "the production runtime… Deterministic, seeded, fully inspectable" (demo.html:14, :149; demo-signal.html:14, :318) BACKED — DEMO-ONLY Engines: mizoki_runtime/demo_*.py (9 desks) — deterministic stdlib fixtures (rollout report note 5) served by the site's production-deployed Flask service (mizoki-website); the pages themselves disclose deterministic/seeded. Site tests: test_demo_platform.py, test_demo_signal.py, per-desk suites. "Production runtime" = the site's own deployed runtime, not platform cells — framing on the page already says so.
C45 Stat row "Native connectors 12 · Governed MCP tools 650+" (index.html:538; grid :537) PARTIAL Of the 12 grid entries, 9 have real in-repo connector code: BigQuery, Neo4j (services/mizoki-journey-kg, graph_writer), Vertex AI (services/vertex-ai-integration), Shopify (intent-shopify-extender, tested), Google Ads (unified_platform + GAQL cell + ekis), Meta Ads (unified_platform + gemini meta worker + ekis), Klaviyo (services/ekis/src/connectors/klaviyo.ts + tests/connectors/klaviyo.test.ts), The Trade Desk (service-marketing-connectors/direct_connectors.py:320-329 TradeDeskAdapter — reporting pull; mutation side is stubs), GitHub (boss/github_actions_integration.py, code_workflow). Gmail / Drive / Calendar have no in-repo API clients — they are external MCP mounts consumed by the Boss surface (boss_agent_core has zero gmail_/gcal_ tools). "650+ tools" is safe-true relative to the repo's own recorded live measurement of 1,137 MCP tools (README.md:348, v6.49.0, 2026-07) — a conservative floor; not re-verified live by this audit.
C46 "CMEK provisioned by Terraform across the reasoning substrate" (index.html:709); "Private VPC and isolated model environments" (:710); privacy.html:405 CMEK statement CMEK: UNBACKED · VPC: PARTIAL Zero CMEK / customer-managed-encryption hits anywhere in deployment/, infrastructure/, gitops/, security/ — the only Terraform KMS is the binary-authorization attestor signing key (deployment/terraform/binary_authorization/), which is not data CMEK. Shared-VPC Terraform is real (deployment/terraform/network.tf modules vpc-prod-shared / vpc-nonprod-shared :2-92; vpc connectors in kg_heartbeat_job) but there is no VPC Service Controls perimeter anywhere. The CMEK sentence on index.html and privacy.html is presently unbacked by IaC in this repo.
C47 "Every proposal, simulation, and authorization is written to a tamper-evident ledger" (index.html:712; :576 "LEDGER APPEND-ONLY · TAMPER-EVIDENT"; :599) BACKED The proven audit chain: contracts/mizoki_contracts/store.py — hash-chained audit records (seq/prev_hash/chain_hash, transactional head doc so "seq/prev_hash cannot race" :18, :174-230) with verify_chain() recomputing the chain (:245-256); TESTED — tests/remediation/test_contracts_hardening.py:239-249 (test_memory_audit_chain_concurrent_appends_stay_intact calls verify_chain); served by services/service-audit-replay/main.py ("Every service appends to the shared hash-chained audit log" :3, verify endpoint via STORE.verify_chain() :58, per-id hash-chained trail :68).

Task-3 cross-cutting observations

  1. The Signal-division dossier numbers are honest to the code: every labeled "operating default" checked (C24 $50, C25 80%, C26 85/15, C27 5%/70%, C28 ±20%/10%, C30 10%, C32 z>2, C33 90/10, C37 Qini 0.02) matches a coded constant exactly. What is missing is tests — the entire boss decisioning family (relu_threshold_agents, autonomous_budget_reallocation_mvp, policy_engine_integration, creative_fatigue_integration, persuadability_segmentation, unified_platform_integration) has zero dedicated unit tests; coverage exists only where mechanisms were reused into cell36 (CUPED, Qini/AUUC gates) or lift-engine (Thompson).
  2. Two demo-page copy defects to fix: C42 (wrong cell number + "served sub-100ms" present-tense design target) and C43 (wrong cell mapping, shipped cell36 omitted). Both are inside a Preview-framed section, but the specific sentences violate the design-target/deployment-claim discipline.
  3. One security-copy defect: C46 — the CMEK-by-Terraform sentence (index.html:709, echoed in privacy.html) has no backing IaC. Either build the Terraform or soften the copy.
  4. DEMO-ONLY verdicts (C21, C44) are accurate as stated on the pages, which disclose deterministic/seeded operation — keep that disclosure intact in any copy edit.

Baseline pin (coordinator directive, build phase concurrent)

All Task 1/2/3 verdicts were measured at 6ef0275 and are hereby pinned to the directed audit baseline 9551ace / e119a1f with no changes required, verified as follows: - 6ef0275 is a git ancestor of both 9551ace and e119a1f (merge-base check). - git diff --name-only 6ef0275 9551ace = only docs/completion-run/{A1_site_content_audit.md, A3_cross_reference_audit.md, PHASE_A_MATRIX.md}; git diff --name-only 6ef0275 e119a1f = only CLAUDE.md, .claude/memory/index.json, and two miz-oki-command-center-ui/ files. Zero overlap with any file cited in any verdict — every file:line citation and quote in this report holds verbatim at both baseline commits. - services/measurement-rails/** is absent from the tree at both 9551ace and e119a1f (git ls-tree) and absent from this container's worktree; no verdict in this report reads or references it. - git status shows no worktree modifications in any cited area (# MIZ OKI 3.5/, services/, src/, miz-oki-adk-agents/, contracts/) in this checkout — no mid-edit fallback to git show HEAD: was needed. - No claim was left UNRESOLVED; the C19–C47 table is complete.

← All docsView source on GitHub →