CC-4 report — Audit pack E: estimand discipline and interval hygiene (WO-10..14, 29, 45)
Date: 2026-09-15 · Branch: audit/cc-4-causal-consolidation · Base: main @ 1f601d171 (2026-09-15)
Lane: MEAS + ENG · Source: docs/audits/AUDIT_EXECUTION_PROMPTS_2026-09-08.md §CC-4, docs/audits/wo/WO-1{0,1,2,3,4}.md, WO-29.md, WO-45.md, docs/audits/AUDIT_2026-09-06_RECONCILIATION.md
PR: to be opened by the coordinator (title: "Audit pack E — estimand discipline and interval hygiene (WO-10..14, 29)"; merging deploys 7 services — see Deploy fan-out)
Claim: not recorded by this session (sandbox: the coordinator records claims)
Every finding was treated as a hypothesis and re-run on 1f601d171 before anything
changed (Step 0). Six of seven reproduced; WO-45 reproduced in part (the v22 pin
had already moved to v23 under WO-24; the default uploadClickConversions send path
and the absence of any refusal still reproduced). Each fix makes a refusal path
explicit, coded and tested; nothing widens access.
Non-independence disclosure (rule 01): the same session wrote every fix and every test here. Step-0 outputs below were captured before the fixes; acceptance runs are this session's own measurements, not an independent verifier's. No live BigQuery, Cloud Run or Google API was touched — every store is an in-process fake with the semantics the WO needs, and the live runs that need credentials are listed under Deferred.
Summary table
| WO | Status | Step-0 evidence on 1f601d171 (before the fix) | Fix, one line |
|---|---|---|---|
| WO-10 | fixed | swapping two purchase timestamps in one cell moved the summed "caused" value $10 ↔ $1,000 (classify() on the audit cell) |
experiment-level estimand (treated mean − control mean) × n_treated with CI; per-purchase split demoted to allocation_convention, label convention, causal=False |
| WO-11 | fixed | Meridian calibration sidecar point_estimate = summed caused value (108.0 in the fixture), standard_error=None; Robyn liftAbs the same |
prior {mean=point, sd=(ci_high−ci_low)/2z} from a WO-10 experiment_estimate ONLY; convention / unlabeled / CI-less input refused by name; export logs experiment id + n |
| WO-12 | fixed | /estimate fit on a 70/30 ROW split and evaluated the policy on the full frame: 84 of 120 rows trained on and evaluated by the same learner |
grouped cross-fitting by user_id, policy on out-of-fold predictions only, CI from a seeded cluster bootstrap over refits, exposure counts on the response |
| WO-13 | fixed | 600 events, P(B>A)=1.0, stored intervals still the prior's [0.025, 0.975], decision CONTINUE | posterior state versioned + append-only; interval recomputed on every update; every reader takes MAX posterior_version per arm |
| WO-14 | fixed | an untagged 1.5× multiplier reached blend_bid_value and produced an applied bid value 73.2 with no provenance anywhere |
multiplier_provenance ∈ {assumption, baseline, measured_effect} + multiplier_source_ref on every effect claim; bid-value and proposal paths accept measured_effect only |
| WO-29 | fixed (fixtures + refusal) | no saturation / carryover fixture; no way to refuse a dose outside support | saturation + carryover synthetic fixtures with known optima; check_dose_request refuses out-of-support doses; constant-dose refusal and the 3.0274 fixture pinned |
| WO-45 | fixed (rail); connector-health row deferred (not CC-4's path) | pin already v23 (WO-24, 2cea2e12f); send_batch still built customers/123:uploadClickConversions by default and nothing refused it |
Data Manager connector request is the default send; legacy builder behind LEGACY_ADS_API_OFFLINE=true refusing past sunset / without pre-2026-06-15 import evidence |
Commits, one per WO, all on 1f601d171:
714d14448 WO-10: Experiment-level estimand replaces individual caused/anticipated labels
17a6ebccb WO-11: Meridian calibration consumes effect + interval only
6e5dcc62d WO-12: Cell 26 honest split and interval
28d9afd2f WO-13: Cell 27 intervals refresh on posterior update
244378719 WO-14: F2 retention multiplier provenance
5808a54d8 WO-29: Dosage estimator nonlinear fixtures and support refusal
804e83c04 WO-45: Retire the uploadClickConversions rail behind the Data Manager path
WO-10 — Experiment-level estimand replaces individual labels
Status: fixed.
Estimand statement (for MEAS to sign). The causal quantity measurement-rails
reports for a registered exposure is the average treatment effect per assignment
unit, scaled to the treated population: ATE_revenue = (mean revenue per treated
unit − mean revenue per holdout unit) × n_treated (and ATE_profit, the same
over per-conversion profit or value × margin_rate), over a declared horizon,
under the holdout assignment named by salt_version (Cell 36's write-once
holdouts:assign surface; the rail never assigns arms). Units that did not convert
contribute 0. The interval is an analytic Welch interval on the difference of unit
means scaled by n_treated (default) or a seeded unit-level percentile bootstrap.
The estimand is a function of per-unit totals only, so it is invariant to the
order, timestamps and count of purchases inside a unit — which is what makes it an
estimand rather than a labeling rule. The per-purchase caused/anticipated split is
retained as a reporting convention: it is emitted only under
allocation_convention with label="convention" on the block and on every ledger
row (causal=False), and it must never feed downstream as causal. The figure is
labeled built, pre-benchmark; nothing here is a performance claim.
Step 0. Pure script on the unmodified module (audit cell: 10 holdout units with one $50 conversion, 10 treated units with a $1,000 and a $10 purchase):
caused sum, big first : 10.0
caused sum, small first: 1000.0
services/measurement-rails/test_wo10_experiment_estimand.py first run on
1f601d171: 11 failed (AttributeError/TypeError: no estimand, no
allocation_convention, horizon/salt_version unknown).
Fix. services/measurement-rails/causal_credit.py: new
estimate_experiment_effect() (schema {label, estimand, point, ci_low, ci_high,
ci_level, ci_method, seed, n_treated, n_control, n_units, treated_mean,
control_mean, horizon, salt_version, experiment_id, estimator_version,
inputs_hash}; refuses unknown estimands, missing horizon/salt, ATE_profit without
profit inputs, more converting units than registered units). classify() returns
{"estimand": …, "allocation_convention": {label, causal=False, method, rows,
baseline_rate, treatment_rate, expected_anticipated, caused, anticipated}, skipped,
written, …}; ledger rows carry label, causal, horizon, salt_version.
main.py /v1/rails/causal-credit:classify requires horizon + salt_version
(400 without) and accepts estimand, margin_rate, ci_method, seed, n_boot,
experiment_id.
Existing 14 causal-credit tests: all kept with explicit successors (they read the
same figures under allocation_convention; docstring says so). API tests updated the
same way; tests/shared/test_origin_strata.py stratifies the convention block (its
rows are still the shipped module's own ledger rows).
Acceptance (11 tests, all green): test_wo10_estimand_is_invariant_to_timestamp_permutation
(point/CI/n identical under the swap; point = $960), test_wo10_step0_convention_caused_sum_moves_under_timestamp_swap
(the convention still moves $10 ↔ $1,000 — kept on the record),
test_wo10_schema_fields_present, test_wo10_convention_output_is_structurally_separate
(no caused/rows at top level; no point/ci in the convention; every row labeled),
test_wo10_ledger_rows_carry_convention_label, test_wo10_horizon_and_salt_version_are_required,
test_wo10_profit_estimand_needs_profit_inputs, test_wo10_unknown_estimand_refused,
test_wo10_more_converting_units_than_registered_units_refused,
test_wo10_seeded_bootstrap_is_reproducible_and_brackets_point,
test_wo10_interval_widens_as_n_shrinks.
Downstream readers of the convention (not changed; reported):
src/shared/growth_control/mmm_export/ledger.py sums classification == caused
(WO-11 now refuses that sum as a calibration input);
miz-oki-adk-agents/boss/causal_credit_ledger_source.py (CX-3's tree) builds ROI
edges from caused rows — it is reading a convention and should consume the
estimand block once its owner picks this up; src/shared/virtuoso_models/
origin_strata.py (dual-homed, fleet deploy) counts rows and accepts a plain row
sequence — callers pass result["allocation_convention"]["rows"]. Marketing claim
A13 ("each conversion classified caused vs anticipated…", canon-locked site) now
describes a convention, not causal evidence — owner wording decision.
Files: services/measurement-rails/{causal_credit.py, main.py, test_causal_credit.py,
test_main_api.py, test_wo10_experiment_estimand.py}, tests/shared/test_origin_strata.py.
WO-11 — Meridian calibration consumes effect + interval only
Status: fixed.
Where the finding lives. The prompt names service-media-incrementality/main.py
and the validation orchestrator. The orchestrator has no Meridian export path
(only the mmm_agreement check over meridian_roi/robyn_roi) — read, nothing to
patch, no patch file needed. The actual R11 site is
src/shared/growth_control/mmm_export/{meridian,robyn}.py (unowned in the table;
src/shared/**, see Deploy fan-out), and service-media-incrementality had no
calibration surface at all.
Step 0. On the unmodified export:
STEP0 meridian calibration: [('google', 108.0, None), ('meta', 108.0, None)] complete= False
tests/governance/test_wo11_meridian_calibration.py first run: collection error
(cannot import name 'calibration').
Fix. New src/shared/growth_control/mmm_export/calibration.py
(calibration_prior(estimate) → {mean, sd, ci_level, experiment_id, n_treated,
n_control, …}; CalibrationRefused with reasons convention_input,
unlabeled_input, missing_ci, degenerate_ci, missing_experiment_id,
missing_n). export_meridian(..., experiment_results={channel: estimate}) and
export_robyn(...) populate the calibration sidecars from these priors only; a
channel without a result carries refusal: no_experiment_result (Meridian) or no
row (Robyn); a convention/CI-less result refuses the whole export; the record logs
experiments: [{channel, experiment_id, n_treated, n_control}]; the caused sum moves
under allocation_convention (label convention). service-media-incrementality
gains POST /api/v1/mmm/calibration-prior (422 with the refusal reason;
unregistered experiments inadmissible; channel must match; audit + log carry
experiment id and n) using a byte-identical twin calibration_prior.py (that image
ships only its own directory + contracts; parity pinned by test).
Acceptance: governance 13 tests green (test_wo11_step0_caused_sum_is_not_the_calibration_point,
test_wo11_valid_experiment_gives_prior_with_positive_sd (sd = 1520/3.92),
test_wo11_convention_input_is_refused, test_wo11_unlabeled_input_is_refused,
test_wo11_input_without_ci_is_refused[ci_low|ci_high],
test_wo11_degenerate_or_infinite_ci_is_refused, test_wo11_experiment_id_and_n_are_required,
test_wo11_meridian_prior_populated_from_experiment_results,
test_wo11_meridian_export_refuses_a_convention_result,
test_wo11_meridian_export_refuses_a_result_without_ci,
test_wo11_meridian_partial_results_are_not_complete,
test_wo11_robyn_calibration_input_uses_the_prior_not_the_caused_sum); service 7 tests
green (tests/remediation/test_wo11_media_incrementality_prior.py, incl. twin parity).
Two pre-existing test_mmm_export.py assertions (point_estimate == 585.0,
liftAbs == "585") replaced by explicit successors; suite 84 green.
Files: src/shared/growth_control/mmm_export/{calibration.py (new), meridian.py,
robyn.py, __init__.py}, services/service-media-incrementality/{main.py,
calibration_prior.py (new)}, tests/governance/{test_wo11_meridian_calibration.py,
test_mmm_export.py}, tests/remediation/test_wo11_media_incrementality_prior.py.
WO-12 — Cell 26: honest split and interval
Status: fixed. Cell 26 = src/cells/cell26/Cell26.py (resolved by CX-1 via
config/actual_urls.py + Dockerfile entrypoint; deploy deploy-cell26-uplift.yml,
dispatch-only).
Step 0. src/cells/cell26/tests/test_wo12_honest_split.py::test_wo12_step0_evaluation_rows_are_disjoint_from_training_rows
first run on 1f601d171:
AssertionError: 84 rows were both trained on and evaluated by the same learner
Fix. Grouped cross-fitting by assignment unit (user_id; a frame without it is
refused 422 — never a silent row split); every row scored out-of-fold; the policy
(AUUC/Qini/value) evaluated on OOF predictions only; the interval is a seeded cluster
bootstrap over refits (resample units, refit the whole cross-fit, dispersion of
the refit means) — < 8 units or degenerate refits → 422 insufficient_clusters;
exposure_counts {rows, treated_rows, control_rows, units, treated_units,
control_units}, cross_fitting and uncertainty blocks on the response.
compute_cate() stays as a compatibility wrapper over estimate_honest().
Reuse, not copy. The fold and bootstrap helpers were extracted from
services/lift-engine/src/core/continuous_dosage.py into
services/lift-engine/src/core/grouped_resampling.py; continuous_dosage.py
imports them (its 18 tests and the 3.0274 fixture unchanged — same RNG consumption),
and Cell26 imports the same module (test_wo12_resampling_utilities_are_imported_from_lift_engine
asserts object identity). The Cell26 Dockerfile copies that one file beside the cell
(build context is the repo root; /services/ is allowlisted in .gcloudignore).
Acceptance (9 tests green): leakage test passes (disjoint), folds grouped by unit,
every row predicted OOF exactly once, exposure counts present,
test_wo12_interval_widens_monotonically_as_n_shrinks (96 → 48 → 24 units, widths
strictly increasing), too-few-units refusal, missing-unit-column refusal, policy from
OOF only, helpers imported. Existing 41 Cell26 tests kept with explicit successors
(fit/predict call shape per fold; SE now from refits — a learner whose predictions do
not depend on the fit gives SE 0 while rows still disperse; a (n, 2) prediction
block is refused instead of averaged); tests/claims_backing/test_a12_cell26_refutation_battery.py
fixture carries user_id and stubs estimate_honest (10 green).
Files: src/cells/cell26/{Cell26.py, Dockerfile, tests/test_cell26_estimators.py,
tests/test_wo12_honest_split.py}, services/lift-engine/src/core/{grouped_resampling.py (new),
continuous_dosage.py}, tests/claims_backing/test_a12_cell26_refutation_battery.py.
WO-13 — Cell 27: intervals refresh on posterior update
Status: fixed. Cell 27 = src/cells/cell27/Cell27.py (the Cell27.py.final* /
_redis_fixed.py siblings are not the Dockerfile entrypoint and were not touched).
Step 0. Throwaway SQL-pattern fake over the unmodified module (applies the exact
UPDATE … SET alpha = alpha + 1 shapes it emits; raises if the interval is ever
written):
posteriors after 600 updates: {'control': {'alpha': 24.0, 'beta': 278.0, 'cr': 0.079, 'ci_lower': 0.025, 'ci_upper': 0.975},
'variant_B': {'alpha': 101.0, 'beta': 201.0, 'cr': 0.334, 'ci_lower': 0.025, 'ci_upper': 0.975}}
decision: CONTINUE | reason: Experiment ongoing | P(B>A)= 1.0
update_posterior never touched credible_interval_95, so promotion compared the
Beta(1,1) prior forever.
Fix. Posterior state is versioned and append-only behind a PosteriorStore
seam: every update inserts a new row for the arm carrying the recomputed
credible_interval_95 and posterior_version = latest + 1; every reader takes the
MAX version per arm (latest_by_arm() in Python; the BigQuery store's SQL is
QUALIFY ROW_NUMBER() OVER (PARTITION BY arm_id ORDER BY posterior_version DESC,
updated_at DESC, row_id DESC) = 1, and the Python rule is applied again on the
result). No DML runs against streamed rows (AGENTS 6.5 — the previous streaming-
insert-then-UPDATE design was also broken on real BigQuery inside the ~90-minute
buffer). Finalization writes a new version per arm; updates on a decided experiment
answer 409. The P(B>A)/futility Monte-Carlo draws are seeded per (experiment,
version) for replay; the redis cache path uses json, not eval.
POST /events/conversion is synchronous and returns the new version + interval.
Schema migration (operator-applied, deferred): posterior_version INT64 and
row_id STRING on unified.online_experiments —
src/cells/cell27/migrations/wo13_posterior_version.sql. Until applied, a streaming
insert with the new columns fails loudly (500), which is the fail-closed direction.
Known limitation, recorded: two concurrent updates to the same arm in the same
instant can both mint v+1 (no CAS on streaming inserts); the reader tie-breaks
deterministically and one increment is lost — same class as the previous design's
lost UPDATE, now visible in the versions.
Acceptance (8 tests green, new offline harness src/cells/cell27/tests/ with
network tripwire): test_wo13_lifecycle_after_n_updates_with_a_clear_winner_promotion_selects_it
(600 events → PROMOTE variant_B, traffic {0, 1}, experiment closed: 404 on read, 409
on update), test_wo13_every_update_recomputes_and_persists_the_interval,
test_wo13_store_is_append_only_never_updated_in_place,
test_wo13_reader_uses_max_version_even_on_shuffled_or_duplicated_rows (stale
duplicate with a newer updated_at loses), test_wo13_bigquery_reader_sql_orders_by_max_version
(+ no UPDATE/DELETE/MERGE in the store), test_wo13_bigquery_store_reads_through_the_client_and_applies_the_rule,
test_wo13_no_winner_is_not_promoted, test_wo13_conversion_endpoint_returns_the_new_version.
Files: src/cells/cell27/{Cell27.py, migrations/wo13_posterior_version.sql (new),
tests/{_env.py, conftest.py, test_wo13_interval_refresh.py} (new)}.
WO-14 — F2 retention multiplier provenance
Status: fixed.
Step 0. On the unmodified package: an untagged [1.5, 1.25] regime evaluated to
a finding whose regime dict had no key mentioning provenance, and
blend_bid_value(..., env={"LTV_BID_WEIGHTING": "true"}) returned
blended_bid_value 73.2, applied=True. tests/governance/test_wo14_f2_multiplier_provenance.py
first run: 15 failed.
Fix. dtr.py: regime keys multiplier_provenance ∈ {assumption, baseline,
measured_effect} + multiplier_source_ref; an effect claim (any multiplier ≠ 1.0)
without them is refused; a no-effect claim defaults to baseline with
observed_curve:<cohort>; baseline on an effect claim is a contradiction (refused);
findings carry provenance per regime and per quarter plus provenance_by_regime and
scenario_only. bid_weighting.blend_bid_value and
plan_integration.horizon_aware_alternatives accept measured_effect only — refused
by name before any record/proposal is emitted, flag or no flag; records and
alternatives carry the provenance and source. dtr.evaluate_regimes (the scenario
path) accepts all three; the arithmetic is provenance-blind by design — only what may
be done with a score differs. services/growth-scheduler/main.py /api/v1/f2/blend
already maps F2BidWeightingError to 422 blend_refused, so the refusal reaches the
service surface without a change there (not CC-4's path; verified by its 58-test suite).
Acceptance (15 tests green): test_wo14_assumption_tagged_multiplier_reaching_bid_value_is_refused,
test_wo14_baseline_tagged_multiplier_reaching_bid_value_is_refused,
test_wo14_measured_effect_reaches_bid_value_with_provenance_on_record,
test_wo14_the_refusal_is_not_flag_gated, test_wo14_assumption_in_the_evaluation_refuses_proposal_generation,
test_wo14_baseline_in_the_evaluation_refuses_proposal_generation,
test_wo14_all_measured_generates_proposals_carrying_provenance,
test_wo14_scenario_path_accepts_all_three_provenances, plus validation tests
(unknown provenance, missing/blank source ref, default baseline, baseline-on-effect
contradiction, scenario_only). tests/governance/test_f2_ltv.py fixtures tagged
(both arms of the holdout experiment; the clamp fixture tagged assumption); 32 + 26
(bid weighting) green. services/net-yield/test_returns_adjusted_f2_bridge.py
(CC-3's) untouched and green — its [1.0] regime is a baseline by default.
Files: src/shared/growth_control/f2_ltv/{dtr.py, bid_weighting.py, plan_integration.py},
tests/governance/{test_f2_ltv.py, test_wo14_f2_multiplier_provenance.py}.
WO-29 — Dosage estimator: nonlinear fixtures and support refusal
Status: fixed (the synthetic half; the bounded prospective real-data test stays under WO-41 / owner).
Fix. services/lift-engine/src/core/continuous_dosage.py: check_dose_request(estimate,
dose) answers supported (tau, se, ci95 at the nearest supported grid point,
requires_approval=True) or insufficient_support / requested_dose_outside_support
with the supported intervals named — below, above, or inside a masked interior gap;
never extrapolates. Shim re-exports it. Fixtures in tests/test_continuous_dosage.py:
simulate_saturation (g(d) = A(1 − e^{−kd}), known d* = −ln(cost/(margin·A·k))/k)
and simulate_carryover (geo × period panel, adstocked dose a_t = d_t + λ d_{t−1},
λ declared by the caller, d* = a*/(1+λ) at steady state).
Measured, honestly. The tau basis is a degree-2 polynomial — a local approximation. On saturation it recovers d* within 5% when doses sit away from the steep near-zero region (fixture support ≈ [1.5, 5.3], n = 2400 rows / 60 groups): errors 0.5 / 4.1 / 4.5 / 3.5 / 0.8 % across seeds 7/13/19/23/29; two seeds (7, 29) are pinned in the test. At n = 600 the same fixture shows 1–16% — sampling variance, and the interval says so. A cubic basis was tried and did not reliably help (4.5–15%) so it was not adopted. Carryover recovered within 10% in adstocked units (4.4% at seed 7) and the raw-dose misspecification moves the optimum (asserted). This is algorithmic evidence on synthetic ground truth; not marketing evidence.
Acceptance (25 tests green, 18 → 25): test_wo29_saturation_optimum_recovered_within_5pct[7|29],
test_wo29_carryover_optimum_recovered_in_adstocked_units_and_mapped_back,
test_wo29_out_of_support_request_is_refused, test_wo29_constant_dose_refusal_is_kept,
test_wo29_existing_linear_fixture_still_recovers_three (d = 3.0274*, pinned < 0.1),
test_wo29_shim_exposes_check_dose_request.
Files: services/lift-engine/{continuous_dosage.py, src/core/continuous_dosage.py,
tests/test_continuous_dosage.py}.
WO-45 — Retire the uploadClickConversions rail behind the Data Manager path
Status: fixed for services/measurement-rails/**; two parts deferred (below).
Step 0 on 1f601d171. GOOGLE_ADS_API_VERSION was already "v23" (WO-24,
2cea2e12f, with the tree-wide sunset guard tests/governance/test_google_ads_api_version_sunset.py
— 3 green on this branch). send_batch still built
customers/{id}:uploadClickConversions by default (the old
test_flag_on_sends_through_injected_transport asserted exactly that) and no flag,
sunset or allowlist check existed. test_wo45_default_path_never_emits_upload_click_conversions
first run: failed on the endpoint.
Fix. Front half unchanged (precedence, 90-day window, before-click — 14 pre-existing
tests green). send_batch(records, customer_id, conversion_action, dry_run=True,
validate_only=True, transport) builds the Data Manager connector's UploadRequest
(/api/v1/upload-conversions → datamanager.googleapis.com/v1/events:ingest) and
never emits an Ads API payload; per-record ad_user_data_consent /
ad_personalization_consent are required (consent_missing); gbraid/wbraid records
are refused by name (data_manager_click_id_unsupported) rather than silently
dropped, because the connector contract carries only gclid — patch for its owner in
docs/audits/patches/WO-45_data_manager_connector_gbraid_wbraid.patch (CX-3, WO-25).
send_batch_legacy_ads_api is compatibility-only: LEGACY_ADS_API_OFFLINE=true
(literal False, source-pinned), refuses api_version_sunset when the pin is past
its recorded sunset (RECORDED_SUNSETS: v21 2026-08-05, v22 2026-10-07, v23 not yet
published), refuses allowlist_evidence_missing / _incomplete / _too_recent without
a recorded import strictly before 2026-06-15; rail flag + transport locks still apply.
Acceptance (14 tests green): default path never emits uploadClickConversions
(dry-run and sent), connector request shape, front-half rules first, consent
required, gbraid/wbraid refused, all-rejected batch refuses, validate_only default,
legacy flag literal, legacy refused flag-off even with evidence, refused at/after
sunset (v22 on 2026-10-07; allowed 2026-10-06), refused without/with incomplete/too
recent evidence, every-lock-open sends the compatibility payload, legacy still needs
the rail flag, the rail's pin is not on the passed-sunset list.
Deferred: (1) validate-only events:ingest against a test account — needs
credentials (shared with WO-25); (2) "Data Manager developer-token / allowlist
eligibility on connector health" lives in services/service-marketing-connectors/
connector_health.py + the UI adapter — unassigned in the ownership table (CX-1 noted
this); not touched. Other uploadClickConversions sites outside CC-4's paths, for the
owner/CX-3: src/cells/cell02/integrations/clients/googleads.client.ts:93,
services/ekis/src/connectors/googleAds/ads.conversions.ts:106,
miz-oki-adk-agents/boss/mcp_connector_registry_v2.py:1025,
miz-oki-adk-agents/boss/unified_platform_integration.py:678.
Files: services/measurement-rails/{offline_conversions.py, test_rails_offline.py,
test_wo45_offline_rail_retirement.py, README.md}, docs/audits/patches/{README.md,
WO-45_data_manager_connector_gbraid_wbraid.patch}.
Gates run (sandbox; no credentials, no network)
| Command | Result |
|---|---|
python3 -m pytest services/measurement-rails -q |
485 passed (459 on base; +26) |
cd services/lift-engine && python3 -m pytest tests/test_continuous_dosage.py -q |
25 passed (18 on base) — tests/test_api.py needs google.cloud.logging, not installable offline; pre-existing |
python3 -m pytest src/cells/cell26/tests -q |
50 passed (41 on base) |
python3 -m pytest src/cells/cell27/tests -q |
8 passed (no harness on base) |
python3 -m pytest tests/governance/test_wo11_meridian_calibration.py tests/governance/test_wo14_f2_multiplier_provenance.py tests/governance/test_mmm_export.py tests/governance/test_f2_ltv.py tests/governance/test_f2_bid_weighting.py -q |
141 passed |
python3 -m pytest tests/governance/test_growth_scheduler.py tests/governance/test_google_ads_api_version_sunset.py -q |
58 + 3 passed |
python3 -m pytest tests/remediation -q --ignore=tests/remediation/test_token_minters.py |
313 passed, 3 failed — the same 3 fail on a pristine 1f601d171 worktree (test_resolve_tenant_*, test_store_failure_degrades_to_env; env-dependent, not this branch's); test_token_minters.py needs google.auth |
python3 -m pytest tests/shared/test_origin_strata.py tests/claims_backing/test_a12_cell26_refutation_battery.py -q |
5 + 10 passed |
python3 -m pytest services/net-yield/test_returns_adjusted_f2_bridge.py -q (CC-3's, read-only) |
4 passed, file untouched |
python3 -m pytest tests/governance -c tests/governance/pytest.ini (whole suite; test_platform_whitepaper_r36.py deselected — needs reportlab, see rule 04) |
1582 passed, 1 failed, 6 skipped — the failure (test_pilot_report.py::test_default_ledger_path_is_the_in_tree_ledger) fails identically on a pristine 1f601d171 worktree (a sibling-worktree path artifact of this sandbox), not this branch's |
bash .github/scripts/content_gates.sh (canon spec self-test + drift ratchet, skill audit, canon/routes/reports/scope/gate-leak/closure-register suites) |
exit 0, 138 passed |
python3 scripts/gate_leak_scan.py --check / tests/reports/test_evidence_manifest.py (both scan this report) |
0 new / 0 grown; 12 passed |
python3 scripts/claude_memory.py check --strict |
memory system is structurally valid (no memory files touched) |
python3 .github/scripts/deploy_router.py --base origin/main --head HEAD |
7 workflows (below) |
Skipped (credentials): live BigQuery for Cell 26/27, Cloud Run health, Data Manager
events:ingest, Meridian/Robyn runs. python3 scripts/claude_memory.py record not run
(coordinator records claims); docs/OPEN_ITEMS.md, CLAUDE.md, .claude/memory/**
untouched.
Deploy fan-out (measured with the router)
deploy-boss-agent-core.yml <- services/measurement-rails/**, src/shared/**
deploy-cell2.yml <- src/shared/**
deploy-cell3.yml <- src/shared/**
deploy-coding-moa.yml <- src/shared/**
deploy-gemini-kg-pipeline.yml <- src/shared/**
deploy-moa-controller.yml <- src/shared/**
deploy-moe-router.yml <- src/shared/**
Merging this PR deploys 7 services (the src/shared/growth_control/** edits fan out
exactly as PR #919 documented). Put the number in the PR title. deploy-gemini-kg-pipeline.yml
is in the set: AGENTS 7.7 — do not merge while a Gemini drain is running.
deploy-cell26-uplift.yml and Cell 27 are dispatch-only and are not dispatched by
this merge; Cell 27 additionally needs the operator migration first.
Files changed (40 files, +3943 / −521)
New: services/measurement-rails/test_wo10_experiment_estimand.py,
services/measurement-rails/test_wo45_offline_rail_retirement.py,
src/shared/growth_control/mmm_export/calibration.py,
services/service-media-incrementality/calibration_prior.py,
tests/governance/test_wo11_meridian_calibration.py,
tests/governance/test_wo14_f2_multiplier_provenance.py,
tests/remediation/test_wo11_media_incrementality_prior.py,
services/lift-engine/src/core/grouped_resampling.py,
src/cells/cell26/tests/test_wo12_honest_split.py,
src/cells/cell27/migrations/wo13_posterior_version.sql,
src/cells/cell27/tests/{_env.py, conftest.py, test_wo13_interval_refresh.py},
docs/audits/patches/{README.md, WO-45_data_manager_connector_gbraid_wbraid.patch},
this report.
Modified: listed per WO above.
Deferred / for the coordinator
- WO-11:
boss/causal_credit_ledger_source.py(CX-3) still reads the convention as ROI evidence — should consume theestimandblock. - WO-13: operator migration
src/cells/cell27/migrations/wo13_posterior_version.sqlbefore the cell is dispatched; live BigQuery run deferred. - WO-29: prospective real-data test under WO-41 (owner).
- WO-45: connector patch (CX-3/WO-25), connector-health eligibility row (unassigned),
validate-only
events:ingestin a test account; four otheruploadClickConversionssites outside CC-4's paths. - Marketing claim A13 wording now describes a convention (canon-locked site; owner).
Proof commands
python3 -m pytest services/measurement-rails -q
( cd services/lift-engine && python3 -m pytest tests/test_continuous_dosage.py -q )
python3 -m pytest src/cells/cell26/tests src/cells/cell27/tests -q
python3 -m pytest tests/governance/test_wo11_meridian_calibration.py tests/governance/test_wo14_f2_multiplier_provenance.py tests/governance/test_mmm_export.py tests/governance/test_f2_ltv.py tests/governance/test_f2_bid_weighting.py tests/governance/test_growth_scheduler.py -q
python3 -m pytest tests/remediation/test_wo11_media_incrementality_prior.py tests/shared/test_origin_strata.py tests/claims_backing/test_a12_cell26_refutation_battery.py services/net-yield/test_returns_adjusted_f2_bridge.py -q
python3 .github/scripts/deploy_router.py --base origin/main --head HEAD
Base SHA the PR is based on: 1f601d171. PR URL: to be opened by the coordinator.