CC-4 report — Audit pack E: estimand discipline and interval hygiene (WO-10..14, 29, 45)

Date: 2026-09-15 · Branch: audit/cc-4-causal-consolidation · Base: main @ 1f601d171 (2026-09-15) Lane: MEAS + ENG · Source: docs/audits/AUDIT_EXECUTION_PROMPTS_2026-09-08.md §CC-4, docs/audits/wo/WO-1{0,1,2,3,4}.md, WO-29.md, WO-45.md, docs/audits/AUDIT_2026-09-06_RECONCILIATION.md PR: to be opened by the coordinator (title: "Audit pack E — estimand discipline and interval hygiene (WO-10..14, 29)"; merging deploys 7 services — see Deploy fan-out) Claim: not recorded by this session (sandbox: the coordinator records claims)

Every finding was treated as a hypothesis and re-run on 1f601d171 before anything changed (Step 0). Six of seven reproduced; WO-45 reproduced in part (the v22 pin had already moved to v23 under WO-24; the default uploadClickConversions send path and the absence of any refusal still reproduced). Each fix makes a refusal path explicit, coded and tested; nothing widens access.

Non-independence disclosure (rule 01): the same session wrote every fix and every test here. Step-0 outputs below were captured before the fixes; acceptance runs are this session's own measurements, not an independent verifier's. No live BigQuery, Cloud Run or Google API was touched — every store is an in-process fake with the semantics the WO needs, and the live runs that need credentials are listed under Deferred.


Summary table

WO Status Step-0 evidence on 1f601d171 (before the fix) Fix, one line
WO-10 fixed swapping two purchase timestamps in one cell moved the summed "caused" value $10 ↔ $1,000 (classify() on the audit cell) experiment-level estimand (treated mean − control mean) × n_treated with CI; per-purchase split demoted to allocation_convention, label convention, causal=False
WO-11 fixed Meridian calibration sidecar point_estimate = summed caused value (108.0 in the fixture), standard_error=None; Robyn liftAbs the same prior {mean=point, sd=(ci_high−ci_low)/2z} from a WO-10 experiment_estimate ONLY; convention / unlabeled / CI-less input refused by name; export logs experiment id + n
WO-12 fixed /estimate fit on a 70/30 ROW split and evaluated the policy on the full frame: 84 of 120 rows trained on and evaluated by the same learner grouped cross-fitting by user_id, policy on out-of-fold predictions only, CI from a seeded cluster bootstrap over refits, exposure counts on the response
WO-13 fixed 600 events, P(B>A)=1.0, stored intervals still the prior's [0.025, 0.975], decision CONTINUE posterior state versioned + append-only; interval recomputed on every update; every reader takes MAX posterior_version per arm
WO-14 fixed an untagged 1.5× multiplier reached blend_bid_value and produced an applied bid value 73.2 with no provenance anywhere multiplier_provenance ∈ {assumption, baseline, measured_effect} + multiplier_source_ref on every effect claim; bid-value and proposal paths accept measured_effect only
WO-29 fixed (fixtures + refusal) no saturation / carryover fixture; no way to refuse a dose outside support saturation + carryover synthetic fixtures with known optima; check_dose_request refuses out-of-support doses; constant-dose refusal and the 3.0274 fixture pinned
WO-45 fixed (rail); connector-health row deferred (not CC-4's path) pin already v23 (WO-24, 2cea2e12f); send_batch still built customers/123:uploadClickConversions by default and nothing refused it Data Manager connector request is the default send; legacy builder behind LEGACY_ADS_API_OFFLINE=true refusing past sunset / without pre-2026-06-15 import evidence

Commits, one per WO, all on 1f601d171:

714d14448 WO-10: Experiment-level estimand replaces individual caused/anticipated labels
17a6ebccb WO-11: Meridian calibration consumes effect + interval only
6e5dcc62d WO-12: Cell 26 honest split and interval
28d9afd2f WO-13: Cell 27 intervals refresh on posterior update
244378719 WO-14: F2 retention multiplier provenance
5808a54d8 WO-29: Dosage estimator nonlinear fixtures and support refusal
804e83c04 WO-45: Retire the uploadClickConversions rail behind the Data Manager path

WO-10 — Experiment-level estimand replaces individual labels

Status: fixed.

Estimand statement (for MEAS to sign). The causal quantity measurement-rails reports for a registered exposure is the average treatment effect per assignment unit, scaled to the treated population: ATE_revenue = (mean revenue per treated unit − mean revenue per holdout unit) × n_treated (and ATE_profit, the same over per-conversion profit or value × margin_rate), over a declared horizon, under the holdout assignment named by salt_version (Cell 36's write-once holdouts:assign surface; the rail never assigns arms). Units that did not convert contribute 0. The interval is an analytic Welch interval on the difference of unit means scaled by n_treated (default) or a seeded unit-level percentile bootstrap. The estimand is a function of per-unit totals only, so it is invariant to the order, timestamps and count of purchases inside a unit — which is what makes it an estimand rather than a labeling rule. The per-purchase caused/anticipated split is retained as a reporting convention: it is emitted only under allocation_convention with label="convention" on the block and on every ledger row (causal=False), and it must never feed downstream as causal. The figure is labeled built, pre-benchmark; nothing here is a performance claim.

Step 0. Pure script on the unmodified module (audit cell: 10 holdout units with one $50 conversion, 10 treated units with a $1,000 and a $10 purchase):

caused sum, big first : 10.0
caused sum, small first: 1000.0

services/measurement-rails/test_wo10_experiment_estimand.py first run on 1f601d171: 11 failed (AttributeError/TypeError: no estimand, no allocation_convention, horizon/salt_version unknown).

Fix. services/measurement-rails/causal_credit.py: new estimate_experiment_effect() (schema {label, estimand, point, ci_low, ci_high, ci_level, ci_method, seed, n_treated, n_control, n_units, treated_mean, control_mean, horizon, salt_version, experiment_id, estimator_version, inputs_hash}; refuses unknown estimands, missing horizon/salt, ATE_profit without profit inputs, more converting units than registered units). classify() returns {"estimand": …, "allocation_convention": {label, causal=False, method, rows, baseline_rate, treatment_rate, expected_anticipated, caused, anticipated}, skipped, written, …}; ledger rows carry label, causal, horizon, salt_version. main.py /v1/rails/causal-credit:classify requires horizon + salt_version (400 without) and accepts estimand, margin_rate, ci_method, seed, n_boot, experiment_id.

Existing 14 causal-credit tests: all kept with explicit successors (they read the same figures under allocation_convention; docstring says so). API tests updated the same way; tests/shared/test_origin_strata.py stratifies the convention block (its rows are still the shipped module's own ledger rows).

Acceptance (11 tests, all green): test_wo10_estimand_is_invariant_to_timestamp_permutation (point/CI/n identical under the swap; point = $960), test_wo10_step0_convention_caused_sum_moves_under_timestamp_swap (the convention still moves $10 ↔ $1,000 — kept on the record), test_wo10_schema_fields_present, test_wo10_convention_output_is_structurally_separate (no caused/rows at top level; no point/ci in the convention; every row labeled), test_wo10_ledger_rows_carry_convention_label, test_wo10_horizon_and_salt_version_are_required, test_wo10_profit_estimand_needs_profit_inputs, test_wo10_unknown_estimand_refused, test_wo10_more_converting_units_than_registered_units_refused, test_wo10_seeded_bootstrap_is_reproducible_and_brackets_point, test_wo10_interval_widens_as_n_shrinks.

Downstream readers of the convention (not changed; reported): src/shared/growth_control/mmm_export/ledger.py sums classification == caused (WO-11 now refuses that sum as a calibration input); miz-oki-adk-agents/boss/causal_credit_ledger_source.py (CX-3's tree) builds ROI edges from caused rows — it is reading a convention and should consume the estimand block once its owner picks this up; src/shared/virtuoso_models/ origin_strata.py (dual-homed, fleet deploy) counts rows and accepts a plain row sequence — callers pass result["allocation_convention"]["rows"]. Marketing claim A13 ("each conversion classified caused vs anticipated…", canon-locked site) now describes a convention, not causal evidence — owner wording decision.

Files: services/measurement-rails/{causal_credit.py, main.py, test_causal_credit.py, test_main_api.py, test_wo10_experiment_estimand.py}, tests/shared/test_origin_strata.py.


WO-11 — Meridian calibration consumes effect + interval only

Status: fixed.

Where the finding lives. The prompt names service-media-incrementality/main.py and the validation orchestrator. The orchestrator has no Meridian export path (only the mmm_agreement check over meridian_roi/robyn_roi) — read, nothing to patch, no patch file needed. The actual R11 site is src/shared/growth_control/mmm_export/{meridian,robyn}.py (unowned in the table; src/shared/**, see Deploy fan-out), and service-media-incrementality had no calibration surface at all.

Step 0. On the unmodified export:

STEP0 meridian calibration: [('google', 108.0, None), ('meta', 108.0, None)] complete= False

tests/governance/test_wo11_meridian_calibration.py first run: collection error (cannot import name 'calibration').

Fix. New src/shared/growth_control/mmm_export/calibration.py (calibration_prior(estimate) → {mean, sd, ci_level, experiment_id, n_treated, n_control, …}; CalibrationRefused with reasons convention_input, unlabeled_input, missing_ci, degenerate_ci, missing_experiment_id, missing_n). export_meridian(..., experiment_results={channel: estimate}) and export_robyn(...) populate the calibration sidecars from these priors only; a channel without a result carries refusal: no_experiment_result (Meridian) or no row (Robyn); a convention/CI-less result refuses the whole export; the record logs experiments: [{channel, experiment_id, n_treated, n_control}]; the caused sum moves under allocation_convention (label convention). service-media-incrementality gains POST /api/v1/mmm/calibration-prior (422 with the refusal reason; unregistered experiments inadmissible; channel must match; audit + log carry experiment id and n) using a byte-identical twin calibration_prior.py (that image ships only its own directory + contracts; parity pinned by test).

Acceptance: governance 13 tests green (test_wo11_step0_caused_sum_is_not_the_calibration_point, test_wo11_valid_experiment_gives_prior_with_positive_sd (sd = 1520/3.92), test_wo11_convention_input_is_refused, test_wo11_unlabeled_input_is_refused, test_wo11_input_without_ci_is_refused[ci_low|ci_high], test_wo11_degenerate_or_infinite_ci_is_refused, test_wo11_experiment_id_and_n_are_required, test_wo11_meridian_prior_populated_from_experiment_results, test_wo11_meridian_export_refuses_a_convention_result, test_wo11_meridian_export_refuses_a_result_without_ci, test_wo11_meridian_partial_results_are_not_complete, test_wo11_robyn_calibration_input_uses_the_prior_not_the_caused_sum); service 7 tests green (tests/remediation/test_wo11_media_incrementality_prior.py, incl. twin parity). Two pre-existing test_mmm_export.py assertions (point_estimate == 585.0, liftAbs == "585") replaced by explicit successors; suite 84 green.

Files: src/shared/growth_control/mmm_export/{calibration.py (new), meridian.py, robyn.py, __init__.py}, services/service-media-incrementality/{main.py, calibration_prior.py (new)}, tests/governance/{test_wo11_meridian_calibration.py, test_mmm_export.py}, tests/remediation/test_wo11_media_incrementality_prior.py.


WO-12 — Cell 26: honest split and interval

Status: fixed. Cell 26 = src/cells/cell26/Cell26.py (resolved by CX-1 via config/actual_urls.py + Dockerfile entrypoint; deploy deploy-cell26-uplift.yml, dispatch-only).

Step 0. src/cells/cell26/tests/test_wo12_honest_split.py::test_wo12_step0_evaluation_rows_are_disjoint_from_training_rows first run on 1f601d171:

AssertionError: 84 rows were both trained on and evaluated by the same learner

Fix. Grouped cross-fitting by assignment unit (user_id; a frame without it is refused 422 — never a silent row split); every row scored out-of-fold; the policy (AUUC/Qini/value) evaluated on OOF predictions only; the interval is a seeded cluster bootstrap over refits (resample units, refit the whole cross-fit, dispersion of the refit means) — < 8 units or degenerate refits → 422 insufficient_clusters; exposure_counts {rows, treated_rows, control_rows, units, treated_units, control_units}, cross_fitting and uncertainty blocks on the response. compute_cate() stays as a compatibility wrapper over estimate_honest().

Reuse, not copy. The fold and bootstrap helpers were extracted from services/lift-engine/src/core/continuous_dosage.py into services/lift-engine/src/core/grouped_resampling.py; continuous_dosage.py imports them (its 18 tests and the 3.0274 fixture unchanged — same RNG consumption), and Cell26 imports the same module (test_wo12_resampling_utilities_are_imported_from_lift_engine asserts object identity). The Cell26 Dockerfile copies that one file beside the cell (build context is the repo root; /services/ is allowlisted in .gcloudignore).

Acceptance (9 tests green): leakage test passes (disjoint), folds grouped by unit, every row predicted OOF exactly once, exposure counts present, test_wo12_interval_widens_monotonically_as_n_shrinks (96 → 48 → 24 units, widths strictly increasing), too-few-units refusal, missing-unit-column refusal, policy from OOF only, helpers imported. Existing 41 Cell26 tests kept with explicit successors (fit/predict call shape per fold; SE now from refits — a learner whose predictions do not depend on the fit gives SE 0 while rows still disperse; a (n, 2) prediction block is refused instead of averaged); tests/claims_backing/test_a12_cell26_refutation_battery.py fixture carries user_id and stubs estimate_honest (10 green).

Files: src/cells/cell26/{Cell26.py, Dockerfile, tests/test_cell26_estimators.py, tests/test_wo12_honest_split.py}, services/lift-engine/src/core/{grouped_resampling.py (new), continuous_dosage.py}, tests/claims_backing/test_a12_cell26_refutation_battery.py.


WO-13 — Cell 27: intervals refresh on posterior update

Status: fixed. Cell 27 = src/cells/cell27/Cell27.py (the Cell27.py.final* / _redis_fixed.py siblings are not the Dockerfile entrypoint and were not touched).

Step 0. Throwaway SQL-pattern fake over the unmodified module (applies the exact UPDATE … SET alpha = alpha + 1 shapes it emits; raises if the interval is ever written):

posteriors after 600 updates: {'control': {'alpha': 24.0, 'beta': 278.0, 'cr': 0.079, 'ci_lower': 0.025, 'ci_upper': 0.975},
                               'variant_B': {'alpha': 101.0, 'beta': 201.0, 'cr': 0.334, 'ci_lower': 0.025, 'ci_upper': 0.975}}
decision: CONTINUE | reason: Experiment ongoing | P(B>A)= 1.0

update_posterior never touched credible_interval_95, so promotion compared the Beta(1,1) prior forever.

Fix. Posterior state is versioned and append-only behind a PosteriorStore seam: every update inserts a new row for the arm carrying the recomputed credible_interval_95 and posterior_version = latest + 1; every reader takes the MAX version per arm (latest_by_arm() in Python; the BigQuery store's SQL is QUALIFY ROW_NUMBER() OVER (PARTITION BY arm_id ORDER BY posterior_version DESC, updated_at DESC, row_id DESC) = 1, and the Python rule is applied again on the result). No DML runs against streamed rows (AGENTS 6.5 — the previous streaming- insert-then-UPDATE design was also broken on real BigQuery inside the ~90-minute buffer). Finalization writes a new version per arm; updates on a decided experiment answer 409. The P(B>A)/futility Monte-Carlo draws are seeded per (experiment, version) for replay; the redis cache path uses json, not eval. POST /events/conversion is synchronous and returns the new version + interval.

Schema migration (operator-applied, deferred): posterior_version INT64 and row_id STRING on unified.online_experiments — src/cells/cell27/migrations/wo13_posterior_version.sql. Until applied, a streaming insert with the new columns fails loudly (500), which is the fail-closed direction. Known limitation, recorded: two concurrent updates to the same arm in the same instant can both mint v+1 (no CAS on streaming inserts); the reader tie-breaks deterministically and one increment is lost — same class as the previous design's lost UPDATE, now visible in the versions.

Acceptance (8 tests green, new offline harness src/cells/cell27/tests/ with network tripwire): test_wo13_lifecycle_after_n_updates_with_a_clear_winner_promotion_selects_it (600 events → PROMOTE variant_B, traffic {0, 1}, experiment closed: 404 on read, 409 on update), test_wo13_every_update_recomputes_and_persists_the_interval, test_wo13_store_is_append_only_never_updated_in_place, test_wo13_reader_uses_max_version_even_on_shuffled_or_duplicated_rows (stale duplicate with a newer updated_at loses), test_wo13_bigquery_reader_sql_orders_by_max_version (+ no UPDATE/DELETE/MERGE in the store), test_wo13_bigquery_store_reads_through_the_client_and_applies_the_rule, test_wo13_no_winner_is_not_promoted, test_wo13_conversion_endpoint_returns_the_new_version.

Files: src/cells/cell27/{Cell27.py, migrations/wo13_posterior_version.sql (new), tests/{_env.py, conftest.py, test_wo13_interval_refresh.py} (new)}.


WO-14 — F2 retention multiplier provenance

Status: fixed.

Step 0. On the unmodified package: an untagged [1.5, 1.25] regime evaluated to a finding whose regime dict had no key mentioning provenance, and blend_bid_value(..., env={"LTV_BID_WEIGHTING": "true"}) returned blended_bid_value 73.2, applied=True. tests/governance/test_wo14_f2_multiplier_provenance.py first run: 15 failed.

Fix. dtr.py: regime keys multiplier_provenance ∈ {assumption, baseline, measured_effect} + multiplier_source_ref; an effect claim (any multiplier ≠ 1.0) without them is refused; a no-effect claim defaults to baseline with observed_curve:<cohort>; baseline on an effect claim is a contradiction (refused); findings carry provenance per regime and per quarter plus provenance_by_regime and scenario_only. bid_weighting.blend_bid_value and plan_integration.horizon_aware_alternatives accept measured_effect only — refused by name before any record/proposal is emitted, flag or no flag; records and alternatives carry the provenance and source. dtr.evaluate_regimes (the scenario path) accepts all three; the arithmetic is provenance-blind by design — only what may be done with a score differs. services/growth-scheduler/main.py /api/v1/f2/blend already maps F2BidWeightingError to 422 blend_refused, so the refusal reaches the service surface without a change there (not CC-4's path; verified by its 58-test suite).

Acceptance (15 tests green): test_wo14_assumption_tagged_multiplier_reaching_bid_value_is_refused, test_wo14_baseline_tagged_multiplier_reaching_bid_value_is_refused, test_wo14_measured_effect_reaches_bid_value_with_provenance_on_record, test_wo14_the_refusal_is_not_flag_gated, test_wo14_assumption_in_the_evaluation_refuses_proposal_generation, test_wo14_baseline_in_the_evaluation_refuses_proposal_generation, test_wo14_all_measured_generates_proposals_carrying_provenance, test_wo14_scenario_path_accepts_all_three_provenances, plus validation tests (unknown provenance, missing/blank source ref, default baseline, baseline-on-effect contradiction, scenario_only). tests/governance/test_f2_ltv.py fixtures tagged (both arms of the holdout experiment; the clamp fixture tagged assumption); 32 + 26 (bid weighting) green. services/net-yield/test_returns_adjusted_f2_bridge.py (CC-3's) untouched and green — its [1.0] regime is a baseline by default.

Files: src/shared/growth_control/f2_ltv/{dtr.py, bid_weighting.py, plan_integration.py}, tests/governance/{test_f2_ltv.py, test_wo14_f2_multiplier_provenance.py}.


WO-29 — Dosage estimator: nonlinear fixtures and support refusal

Status: fixed (the synthetic half; the bounded prospective real-data test stays under WO-41 / owner).

Fix. services/lift-engine/src/core/continuous_dosage.py: check_dose_request(estimate, dose) answers supported (tau, se, ci95 at the nearest supported grid point, requires_approval=True) or insufficient_support / requested_dose_outside_support with the supported intervals named — below, above, or inside a masked interior gap; never extrapolates. Shim re-exports it. Fixtures in tests/test_continuous_dosage.py: simulate_saturation (g(d) = A(1 − e^{−kd}), known d* = −ln(cost/(margin·A·k))/k) and simulate_carryover (geo × period panel, adstocked dose a_t = d_t + λ d_{t−1}, λ declared by the caller, d* = a*/(1+λ) at steady state).

Measured, honestly. The tau basis is a degree-2 polynomial — a local approximation. On saturation it recovers d* within 5% when doses sit away from the steep near-zero region (fixture support ≈ [1.5, 5.3], n = 2400 rows / 60 groups): errors 0.5 / 4.1 / 4.5 / 3.5 / 0.8 % across seeds 7/13/19/23/29; two seeds (7, 29) are pinned in the test. At n = 600 the same fixture shows 1–16% — sampling variance, and the interval says so. A cubic basis was tried and did not reliably help (4.5–15%) so it was not adopted. Carryover recovered within 10% in adstocked units (4.4% at seed 7) and the raw-dose misspecification moves the optimum (asserted). This is algorithmic evidence on synthetic ground truth; not marketing evidence.

Acceptance (25 tests green, 18 → 25): test_wo29_saturation_optimum_recovered_within_5pct[7|29], test_wo29_carryover_optimum_recovered_in_adstocked_units_and_mapped_back, test_wo29_out_of_support_request_is_refused, test_wo29_constant_dose_refusal_is_kept, test_wo29_existing_linear_fixture_still_recovers_three (d = 3.0274*, pinned < 0.1), test_wo29_shim_exposes_check_dose_request.

Files: services/lift-engine/{continuous_dosage.py, src/core/continuous_dosage.py, tests/test_continuous_dosage.py}.


WO-45 — Retire the uploadClickConversions rail behind the Data Manager path

Status: fixed for services/measurement-rails/**; two parts deferred (below).

Step 0 on 1f601d171. GOOGLE_ADS_API_VERSION was already "v23" (WO-24, 2cea2e12f, with the tree-wide sunset guard tests/governance/test_google_ads_api_version_sunset.py — 3 green on this branch). send_batch still built customers/{id}:uploadClickConversions by default (the old test_flag_on_sends_through_injected_transport asserted exactly that) and no flag, sunset or allowlist check existed. test_wo45_default_path_never_emits_upload_click_conversions first run: failed on the endpoint.

Fix. Front half unchanged (precedence, 90-day window, before-click — 14 pre-existing tests green). send_batch(records, customer_id, conversion_action, dry_run=True, validate_only=True, transport) builds the Data Manager connector's UploadRequest (/api/v1/upload-conversions → datamanager.googleapis.com/v1/events:ingest) and never emits an Ads API payload; per-record ad_user_data_consent / ad_personalization_consent are required (consent_missing); gbraid/wbraid records are refused by name (data_manager_click_id_unsupported) rather than silently dropped, because the connector contract carries only gclid — patch for its owner in docs/audits/patches/WO-45_data_manager_connector_gbraid_wbraid.patch (CX-3, WO-25). send_batch_legacy_ads_api is compatibility-only: LEGACY_ADS_API_OFFLINE=true (literal False, source-pinned), refuses api_version_sunset when the pin is past its recorded sunset (RECORDED_SUNSETS: v21 2026-08-05, v22 2026-10-07, v23 not yet published), refuses allowlist_evidence_missing / _incomplete / _too_recent without a recorded import strictly before 2026-06-15; rail flag + transport locks still apply.

Acceptance (14 tests green): default path never emits uploadClickConversions (dry-run and sent), connector request shape, front-half rules first, consent required, gbraid/wbraid refused, all-rejected batch refuses, validate_only default, legacy flag literal, legacy refused flag-off even with evidence, refused at/after sunset (v22 on 2026-10-07; allowed 2026-10-06), refused without/with incomplete/too recent evidence, every-lock-open sends the compatibility payload, legacy still needs the rail flag, the rail's pin is not on the passed-sunset list.

Deferred: (1) validate-only events:ingest against a test account — needs credentials (shared with WO-25); (2) "Data Manager developer-token / allowlist eligibility on connector health" lives in services/service-marketing-connectors/ connector_health.py + the UI adapter — unassigned in the ownership table (CX-1 noted this); not touched. Other uploadClickConversions sites outside CC-4's paths, for the owner/CX-3: src/cells/cell02/integrations/clients/googleads.client.ts:93, services/ekis/src/connectors/googleAds/ads.conversions.ts:106, miz-oki-adk-agents/boss/mcp_connector_registry_v2.py:1025, miz-oki-adk-agents/boss/unified_platform_integration.py:678.

Files: services/measurement-rails/{offline_conversions.py, test_rails_offline.py, test_wo45_offline_rail_retirement.py, README.md}, docs/audits/patches/{README.md, WO-45_data_manager_connector_gbraid_wbraid.patch}.


Gates run (sandbox; no credentials, no network)

Command Result
python3 -m pytest services/measurement-rails -q 485 passed (459 on base; +26)
cd services/lift-engine && python3 -m pytest tests/test_continuous_dosage.py -q 25 passed (18 on base) — tests/test_api.py needs google.cloud.logging, not installable offline; pre-existing
python3 -m pytest src/cells/cell26/tests -q 50 passed (41 on base)
python3 -m pytest src/cells/cell27/tests -q 8 passed (no harness on base)
python3 -m pytest tests/governance/test_wo11_meridian_calibration.py tests/governance/test_wo14_f2_multiplier_provenance.py tests/governance/test_mmm_export.py tests/governance/test_f2_ltv.py tests/governance/test_f2_bid_weighting.py -q 141 passed
python3 -m pytest tests/governance/test_growth_scheduler.py tests/governance/test_google_ads_api_version_sunset.py -q 58 + 3 passed
python3 -m pytest tests/remediation -q --ignore=tests/remediation/test_token_minters.py 313 passed, 3 failed — the same 3 fail on a pristine 1f601d171 worktree (test_resolve_tenant_*, test_store_failure_degrades_to_env; env-dependent, not this branch's); test_token_minters.py needs google.auth
python3 -m pytest tests/shared/test_origin_strata.py tests/claims_backing/test_a12_cell26_refutation_battery.py -q 5 + 10 passed
python3 -m pytest services/net-yield/test_returns_adjusted_f2_bridge.py -q (CC-3's, read-only) 4 passed, file untouched
python3 -m pytest tests/governance -c tests/governance/pytest.ini (whole suite; test_platform_whitepaper_r36.py deselected — needs reportlab, see rule 04) 1582 passed, 1 failed, 6 skipped — the failure (test_pilot_report.py::test_default_ledger_path_is_the_in_tree_ledger) fails identically on a pristine 1f601d171 worktree (a sibling-worktree path artifact of this sandbox), not this branch's
bash .github/scripts/content_gates.sh (canon spec self-test + drift ratchet, skill audit, canon/routes/reports/scope/gate-leak/closure-register suites) exit 0, 138 passed
python3 scripts/gate_leak_scan.py --check / tests/reports/test_evidence_manifest.py (both scan this report) 0 new / 0 grown; 12 passed
python3 scripts/claude_memory.py check --strict memory system is structurally valid (no memory files touched)
python3 .github/scripts/deploy_router.py --base origin/main --head HEAD 7 workflows (below)

Skipped (credentials): live BigQuery for Cell 26/27, Cloud Run health, Data Manager events:ingest, Meridian/Robyn runs. python3 scripts/claude_memory.py record not run (coordinator records claims); docs/OPEN_ITEMS.md, CLAUDE.md, .claude/memory/** untouched.

Deploy fan-out (measured with the router)

deploy-boss-agent-core.yml  <- services/measurement-rails/**, src/shared/**
deploy-cell2.yml            <- src/shared/**
deploy-cell3.yml            <- src/shared/**
deploy-coding-moa.yml       <- src/shared/**
deploy-gemini-kg-pipeline.yml <- src/shared/**
deploy-moa-controller.yml   <- src/shared/**
deploy-moe-router.yml       <- src/shared/**

Merging this PR deploys 7 services (the src/shared/growth_control/** edits fan out exactly as PR #919 documented). Put the number in the PR title. deploy-gemini-kg-pipeline.yml is in the set: AGENTS 7.7 — do not merge while a Gemini drain is running. deploy-cell26-uplift.yml and Cell 27 are dispatch-only and are not dispatched by this merge; Cell 27 additionally needs the operator migration first.

Files changed (40 files, +3943 / −521)

New: services/measurement-rails/test_wo10_experiment_estimand.py, services/measurement-rails/test_wo45_offline_rail_retirement.py, src/shared/growth_control/mmm_export/calibration.py, services/service-media-incrementality/calibration_prior.py, tests/governance/test_wo11_meridian_calibration.py, tests/governance/test_wo14_f2_multiplier_provenance.py, tests/remediation/test_wo11_media_incrementality_prior.py, services/lift-engine/src/core/grouped_resampling.py, src/cells/cell26/tests/test_wo12_honest_split.py, src/cells/cell27/migrations/wo13_posterior_version.sql, src/cells/cell27/tests/{_env.py, conftest.py, test_wo13_interval_refresh.py}, docs/audits/patches/{README.md, WO-45_data_manager_connector_gbraid_wbraid.patch}, this report. Modified: listed per WO above.

Deferred / for the coordinator

Proof commands

python3 -m pytest services/measurement-rails -q
( cd services/lift-engine && python3 -m pytest tests/test_continuous_dosage.py -q )
python3 -m pytest src/cells/cell26/tests src/cells/cell27/tests -q
python3 -m pytest tests/governance/test_wo11_meridian_calibration.py tests/governance/test_wo14_f2_multiplier_provenance.py tests/governance/test_mmm_export.py tests/governance/test_f2_ltv.py tests/governance/test_f2_bid_weighting.py tests/governance/test_growth_scheduler.py -q
python3 -m pytest tests/remediation/test_wo11_media_incrementality_prior.py tests/shared/test_origin_strata.py tests/claims_backing/test_a12_cell26_refutation_battery.py services/net-yield/test_returns_adjusted_f2_bridge.py -q
python3 .github/scripts/deploy_router.py --base origin/main --head HEAD

Base SHA the PR is based on: 1f601d171. PR URL: to be opened by the coordinator.

← All docsView source on GitHub →