ADR-MR-001 — Measurement Rails service

Status: accepted (scaffold) · Date: 2026-08-08 · claim_label: built, pre-benchmark Deploy state: implemented — NOT deployed. There is no push-triggered deploy workflow for services/measurement-rails/**; the only path to a serving revision is the operator RUNBOOK (docs/measurement-rails/RUNBOOK.md).

Context

The platform needs one measurement backbone that (a) re-attributes every conversion under a single house rule instead of trusting each platform's self-graded homework, (b) turns holdout membership into an explicit caused/anticipated ledger, (c) watches platform-vs-house drift, and (d) prepares — without autonomously firing — corrected-value writebacks to Google/Meta/GA4. The owner approved this build on the completion-run fix list. Everything ships flag-off, dry-run, and router-inert.

Decisions

1. New router-inert service, with PROOF

A new bounded service services/measurement-rails/ (net-yield's pattern: no deploy-*.yml, cloudbuild is build+push only, RUNBOOK-only deploys). Router-match proof — every .github/workflows/*.yml on.push.paths pattern was matched (GitHub glob semantics: ** crosses /, * does not) against every created path (services/measurement-rails/**, bigquery/schemas/measurement_rails_ddl.sql, config/measurement_rails/**, docs/measurement-rails/**). Result, verbatim (2026-08-08, branch completion-run):

== push-path MATCHES against measurement-rails owned paths ==
.github/workflows/validate-manifests.yml: pattern 'services/**/cloudbuild.yaml' matches services/measurement-rails/cloudbuild.yaml

== workflows with push triggers and NO path filter (fire on any push to their branches) ==
.github/workflows/auto-merge-ai-branches.yml: branches=['claude/**', 'gemini/**', 'codex/**', 'copilot/**', 'cursor/**']
.github/workflows/auto-merge-grothendieck.yml: branches=['romantic-grothendieck']
.github/workflows/auto-merge-mcclintock.yml: branches=['beautiful-mcclintock']
.github/workflows/deploy-gcs.yml: branches=['**']
.github/workflows/verify-no-secrets.yml: branches=['main', 'develop']

Reading: ZERO deploy workflows match any created path. The single path-filtered match is validate-manifests.yml — a read-only lint (permissions: contents: read) that validates cloudbuild manifests and deploys nothing; scripts/validate_manifests.py services/measurement-rails/cloudbuild.yaml passes ("All validations passed"). The path-independent workflows (repo backup to GCS, secret scan, bot-branch auto-merge) fire on ANY push to their branches regardless of which files change — they are not path matches and deploy no service.

2. Supersedes — does not edit — the legacy boss mock paths

Amendment 2026-08-09 — one line of the boss WAS changed, for a live-send hole. miz-oki-adk-agents/boss/enhanced_conversions_integration.py is not a mock path: it holds a real httpx.AsyncClient, a grep for dry_run|DRY_RUN|dry-run across it returns zero hits, and boss_agent_core.py:1803 gated it behind os.getenv("ENABLE_ENHANCED_CONVERSIONS", "true") — a default-ON flag for a module with no dry-run mode. Only empty default credentials stood between that and live writes to Google/Meta, which contradicts the rule that all external-platform writeback ships flag-gated OFF.

The default is now "false". That is the entire change: one line, no edit to enhanced_conversions_integration.py, nothing else in boss_agent_core.py (which is deploy-hot — .github/workflows/deploy-boss-agent-core.yml triggers on that exact path). The flag guards initialization plus endpoint and MCP-tool registration, and its false branch is the same one already taken when ENHANCED_CONVERSIONS_AVAILABLE is false, so the skip is an existing, exercised path. Operators who want the integration set ENABLE_ENHANCED_CONVERSIONS=true explicitly — an affirmative act, which is what enabling an undryrunnable live sender should require. Retiring or dry-run-arming that module remains a boss-image migration item, unchanged by this.

miz-oki-adk-agents/boss/cross_platform_attribution_integration.py and conversion_tracking_integration.py contain mocked data-fetch and mocked API-execution paths ("Mocking data fetch", "Execute API Call (Mocked/Placeholder)", mock-mode service markers). This service supersedes those capabilities with tested, flag-gated, transport-injected implementations. The legacy modules are NOT edited or deleted here: the boss image vendors its modules (AGENTS.md Article 5.4 — grep Dockerfiles before touching), so their retirement is a RECORDED MIGRATION ITEM for a boss-image release: (1) point boss tools at the measurement-rails API, (2) verify zero importers, (3) delete in a boss-image commit with its own deploy verification. Never a bare delete from this path.

3. Dataset reality: mizoki_unified_data, not unified

The spec named unified.house_attribution / unified.causal_credit_ledger. The platform's revenue dataset is mizoki_unified_data and the spec itself directs matching net-yield's dataset resolution (BQ_DATASET env, default mizoki_unified_data — ADR-NY-001 §4 precedent: reality wins). Tables: mizoki_unified_data.house_attribution, mizoki_unified_data.causal_credit_ledger, partitioned by DATE(occurred_at) and clustered per the existing conventions, with *_latest views mirroring net_contribution_cohort_latest.

4. Two hash families — platform match keys vs internal keys

Google Enhanced Conversions and Meta CAPI match on UNSALTED SHA-256 of the platform-normalized identifier; salting outbound hashes would silently zero the match rate. So identity.platform_hash is unsalted-normalized (Google: trim/lowercase/gmail-dot-removal, E.164 phones; Meta: trim/lowercase, digits-only phones) and appears ONLY in outbound payloads (which ship nowhere by default), while identity.internal_key — peppered SHA-256 (MEASUREMENT_IDENTITY_PEPPER + fixed context string) — is the only hash family for anything this platform persists or returns. The spec's "salted SHA-256" requirement is honored where salting is sound (our persistence) and documented where the platform spec forbids it (their match keys). Raw identifiers never persist, never log; tests scan artifacts for the raw fixtures.

5. Deterministic-only identity for causal math

identity.deterministic_only rejects probabilistic identities from causal_credit outright (AGENTS.md Article 6.6; Cell 36 precedent — probabilistic identities are rejected, never down-weighted). Units without registered holdout membership are SKIPPED with a named reason — membership authority stays with Cell 36's write-once holdouts:assign surface; this service composes, it never assigns arms.

Amendment (PR #644 review): "composes, never assigns" was true of the code and still let the ledger be fabricated. holdout_assignments arrived in the REQUEST BODY, so a caller could hand in any arms it liked along with the treatment_unit_count / holdout_unit_count that set the baseline — and a caller who sets the baseline sets how many rows come back caused. Composing faithfully from an input nobody verified is not a safety property.

The decision, therefore, is about PROVENANCE rather than assignment: causal-credit:classify resolves arms from a server-side app.state.holdout_lookup when one is wired, and only such a classification may be PERSISTED. A body-supplied map still classifies — labelled arm_source: "caller_supplied_unverified" on the response — and a persist request on it is refused 403 unverified_arms. The refusal is not flag-gated: MEASUREMENT_PERSIST states whether writes are allowed, never whether the input deserved to be written.

This deliberately does NOT introduce a Cell 36 HTTP client here — §6's no-HTTP-client property is load-bearing and stays. The lookup is injected at composition time (RUNBOOK §5c), which also means it is currently unwired and the route is a dry-run surface until an operator binds it.

6. Transport-injection dry-run design

Every outbound-capable module takes an injected transport callable and constructs no HTTP client of its own. A live send requires three things at once: the rail's flag on (flags.require_rail), dry_run=False per call, and an injected transport (flags.require_transport). Tests therefore cannot reach the network by construction, and the shipped state (all flags false, dry_run defaults true, no transports wired) is inert three ways. GA4 goes further: send() only accepts a ValidatedPayload that only validate() can construct — send-without-validate is a TypeError.

7. 48h Meta dedup window to spec

meta_capi.DEDUP_WINDOW_HOURS = 48 matches Meta's documented event_id/event_name dedup horizon for pixel/CAPI pairs; the shared id shape is mr-{tenant}-{source_event_id}. A unit test asserts the literal 48 so the window cannot drift silently; boundary tests cover 47h (duplicate), 48h and 49h (new).

7a. Persistence is a third lock, not a code path (amended 2026-08-09)

Amends §6's scope. As first built, the only recompute surface was POST /v1/rails/recompute:dry-run against an in-memory writer: MeasurementRailsBQWriter existed and was tested but was instantiated nowhere in production code, so no persisted row could ever be written and causal_credit had no HTTP caller at all. The tables and *_latest views were real; nothing could fill them.

The wiring adds POST /v1/rails/recompute and POST /v1/rails/causal-credit:classify WITHOUT touching the attribution or classification logic (same recompute(), same 7d-click/1d-view window, same classify() with its deterministic-only guard) — the routes differ from the dry-run route only in which writer they are handed.

Persistence borrows §6's shape rather than inventing a new one: a live platform send needs flag + dry_run=False + injected transport, so a persisted write needs MEASUREMENT_PERSIST=true + per-call "persist": true + a configured writer. Rejected alternative: making persistence the default whenever GCP_PROJECT happens to be set — deployment configuration is not consent to write, and it would make a misrouted scheduler job a data-writing event.

A half-open configuration REFUSES (403 flag off, 503 no writer, 502 insert failure) instead of falling back to the in-memory writer. A caller that asked to persist and received "written": 1 from an in-memory writer would have been told rows landed in BigQuery when they did not.

MEASUREMENT_PERSIST is deliberately NOT in the §6 outbound ladder and is kept semantically distinct from MEASUREMENT_WRITEBACK: it governs first-party storage in the tenant's own dataset and can never cause an outbound platform send. It is reported in flag_states() (so /health and /metrics show it) and its False default is pinned as source text by test_flags.py, like the other defaults; the ambient sentinel now also turns the suite red if the suite is invoked with it enabled.

7b. Drift streak state is durable, and says when it is not (2026-08-09)

The drift monitor held streaks in a process dict, so a Cloud Run instance recycling on day 2 of a 3-day drift silently reset the count toward the ALERT_CONSECUTIVE_DAYS = 3 threshold. An alert that never fires is indistinguishable from "no drift" — a fabricated measurement by omission.

An injected observation store (bq.DriftObservationStore, the same append-only injected-client pattern as MeasurementRailsBQWriter, plus the read half needed to rehydrate) makes record_day append and hydrate() rebuild the streak at construction. Re-recording a day appends; the read keeps the newest row per (tenant_id, source, day) by recorded_at — latest-wins, mirroring the *_latest views. No UPDATE path was added.

The store stays OPTIONAL because it needs operator-applied DDL (RUNBOOK §1a) and credentials, so the honest posture is to report the difference rather than imply durability: status() and /health carry durable / drift_state_durable, false when the streak is memory-only. A store append that fails RAISES (502 from the route) — a day the caller believes was recorded but that no restart can recover is silent data loss.

write_alerts had an injected sink but no caller; the record route now invokes it, defaulting to a structured Cloud Logging line. That default was chosen over a webhook because it requires no credential and no HTTP client. The constraint that must keep holding is a property of the CODE: no measurement-rails module imports or constructs an HTTP client — enforced by test_flags.py::TestNoHttpClientInRailsModules, which scans every non-test module in the service directory.

Correction (recorded rather than quietly edited): an earlier revision of this section stated the constraint as "requirements.txt contains no HTTP client at all". That is false. requirements.txt names no HTTP library directly, but the pinned google-cloud-bigquery==3.27.0 pulls requests and urllib3 into the image transitively, so the image DOES ship an HTTP stack. The behavioural property was and is true; the stated basis for it was not. Dependency absence was never the real guarantee anyway — §6's transport injection is: a send needs a flag on, dry_run=False, and an explicitly injected transport, none of which an installed library provides.

Consequence: measurement_drift_observations is a new table whose DDL currently lives in RUNBOOK §1a rather than bigquery/schemas/measurement_rails_ddl.sql. Folding it in is a tracked follow-up (RUNBOOK §8).

8. Bitemporal honesty is structural

house_attribution and causal_credit stamp available_to_model_at at recompute/classification time, never business time, and RAISE BitemporalViolation when it would precede observed_at (Article 6.1). The causal ledger is append-only — the service has no UPDATE path (Article 6.2).

9. Registry before exposure

design_registry requires mde up front, bounds holdout_pct to [0.10, 0.20] for intent-driven activation designs (Cell 36 ships 10% permanent holdout), enforces write-once (409 on re-register), and refuses designs whose exposure already started — pre-registration that can be back-dated is not pre-registration.

Consequences

← All docsView source on GitHub →