ADR-MR-001 — Measurement Rails service
Status: accepted (scaffold) · Date: 2026-08-08 · claim_label: built, pre-benchmark
Deploy state: implemented — NOT deployed. There is no push-triggered
deploy workflow for services/measurement-rails/**; the only path to a
serving revision is the operator RUNBOOK (docs/measurement-rails/RUNBOOK.md).
Context
The platform needs one measurement backbone that (a) re-attributes every conversion under a single house rule instead of trusting each platform's self-graded homework, (b) turns holdout membership into an explicit caused/anticipated ledger, (c) watches platform-vs-house drift, and (d) prepares — without autonomously firing — corrected-value writebacks to Google/Meta/GA4. The owner approved this build on the completion-run fix list. Everything ships flag-off, dry-run, and router-inert.
Decisions
1. New router-inert service, with PROOF
A new bounded service services/measurement-rails/ (net-yield's pattern:
no deploy-*.yml, cloudbuild is build+push only, RUNBOOK-only deploys).
Router-match proof — every .github/workflows/*.yml on.push.paths
pattern was matched (GitHub glob semantics: ** crosses /, * does not)
against every created path
(services/measurement-rails/**, bigquery/schemas/measurement_rails_ddl.sql,
config/measurement_rails/**, docs/measurement-rails/**). Result,
verbatim (2026-08-08, branch completion-run):
== push-path MATCHES against measurement-rails owned paths ==
.github/workflows/validate-manifests.yml: pattern 'services/**/cloudbuild.yaml' matches services/measurement-rails/cloudbuild.yaml
== workflows with push triggers and NO path filter (fire on any push to their branches) ==
.github/workflows/auto-merge-ai-branches.yml: branches=['claude/**', 'gemini/**', 'codex/**', 'copilot/**', 'cursor/**']
.github/workflows/auto-merge-grothendieck.yml: branches=['romantic-grothendieck']
.github/workflows/auto-merge-mcclintock.yml: branches=['beautiful-mcclintock']
.github/workflows/deploy-gcs.yml: branches=['**']
.github/workflows/verify-no-secrets.yml: branches=['main', 'develop']
Reading: ZERO deploy workflows match any created path. The single
path-filtered match is validate-manifests.yml — a read-only lint
(permissions: contents: read) that validates cloudbuild manifests and
deploys nothing; scripts/validate_manifests.py
services/measurement-rails/cloudbuild.yaml passes ("All validations
passed"). The path-independent workflows (repo backup to GCS, secret scan,
bot-branch auto-merge) fire on ANY push to their branches regardless of
which files change — they are not path matches and deploy no service.
2. Supersedes — does not edit — the legacy boss mock paths
Amendment 2026-08-09 — one line of the boss WAS changed, for a live-send
hole. miz-oki-adk-agents/boss/enhanced_conversions_integration.py is not
a mock path: it holds a real httpx.AsyncClient, a grep for
dry_run|DRY_RUN|dry-run across it returns zero hits, and
boss_agent_core.py:1803 gated it behind
os.getenv("ENABLE_ENHANCED_CONVERSIONS", "true") — a default-ON flag for a
module with no dry-run mode. Only empty default credentials stood between
that and live writes to Google/Meta, which contradicts the rule that all
external-platform writeback ships flag-gated OFF.
The default is now "false". That is the entire change: one line, no edit to
enhanced_conversions_integration.py, nothing else in boss_agent_core.py
(which is deploy-hot — .github/workflows/deploy-boss-agent-core.yml
triggers on that exact path). The flag guards initialization plus endpoint
and MCP-tool registration, and its false branch is the same one already
taken when ENHANCED_CONVERSIONS_AVAILABLE is false, so the skip is an
existing, exercised path. Operators who want the integration set
ENABLE_ENHANCED_CONVERSIONS=true explicitly — an affirmative act, which is
what enabling an undryrunnable live sender should require. Retiring or
dry-run-arming that module remains a boss-image migration item, unchanged by
this.
miz-oki-adk-agents/boss/cross_platform_attribution_integration.py and
conversion_tracking_integration.py contain mocked data-fetch and mocked
API-execution paths ("Mocking data fetch", "Execute API Call
(Mocked/Placeholder)", mock-mode service markers). This service supersedes
those capabilities with tested, flag-gated, transport-injected
implementations. The legacy modules are NOT edited or deleted here: the
boss image vendors its modules (AGENTS.md Article 5.4 — grep Dockerfiles
before touching), so their retirement is a RECORDED MIGRATION ITEM for a
boss-image release: (1) point boss tools at the measurement-rails API,
(2) verify zero importers, (3) delete in a boss-image commit with its own
deploy verification. Never a bare delete from this path.
3. Dataset reality: mizoki_unified_data, not unified
The spec named unified.house_attribution / unified.causal_credit_ledger.
The platform's revenue dataset is mizoki_unified_data and the spec itself
directs matching net-yield's dataset resolution (BQ_DATASET env, default
mizoki_unified_data — ADR-NY-001 §4 precedent: reality wins). Tables:
mizoki_unified_data.house_attribution,
mizoki_unified_data.causal_credit_ledger, partitioned by
DATE(occurred_at) and clustered per the existing conventions, with
*_latest views mirroring net_contribution_cohort_latest.
4. Two hash families — platform match keys vs internal keys
Google Enhanced Conversions and Meta CAPI match on UNSALTED SHA-256 of the
platform-normalized identifier; salting outbound hashes would silently zero
the match rate. So identity.platform_hash is unsalted-normalized
(Google: trim/lowercase/gmail-dot-removal, E.164 phones; Meta:
trim/lowercase, digits-only phones) and appears ONLY in outbound payloads
(which ship nowhere by default), while identity.internal_key — peppered
SHA-256 (MEASUREMENT_IDENTITY_PEPPER + fixed context string) — is the
only hash family for anything this platform persists or returns. The spec's
"salted SHA-256" requirement is honored where salting is sound (our
persistence) and documented where the platform spec forbids it (their match
keys). Raw identifiers never persist, never log; tests scan artifacts for
the raw fixtures.
5. Deterministic-only identity for causal math
identity.deterministic_only rejects probabilistic identities from
causal_credit outright (AGENTS.md Article 6.6; Cell 36 precedent —
probabilistic identities are rejected, never down-weighted). Units without
registered holdout membership are SKIPPED with a named reason — membership
authority stays with Cell 36's write-once holdouts:assign surface; this
service composes, it never assigns arms.
Amendment (PR #644 review): "composes, never assigns" was true of the
code and still let the ledger be fabricated. holdout_assignments arrived in
the REQUEST BODY, so a caller could hand in any arms it liked along with the
treatment_unit_count / holdout_unit_count that set the baseline — and a
caller who sets the baseline sets how many rows come back caused. Composing
faithfully from an input nobody verified is not a safety property.
The decision, therefore, is about PROVENANCE rather than assignment:
causal-credit:classify resolves arms from a server-side
app.state.holdout_lookup when one is wired, and only such a classification
may be PERSISTED. A body-supplied map still classifies — labelled
arm_source: "caller_supplied_unverified" on the response — and a persist
request on it is refused 403 unverified_arms. The refusal is not
flag-gated: MEASUREMENT_PERSIST states whether writes are allowed, never
whether the input deserved to be written.
This deliberately does NOT introduce a Cell 36 HTTP client here — §6's no-HTTP-client property is load-bearing and stays. The lookup is injected at composition time (RUNBOOK §5c), which also means it is currently unwired and the route is a dry-run surface until an operator binds it.
6. Transport-injection dry-run design
Every outbound-capable module takes an injected transport callable and
constructs no HTTP client of its own. A live send requires three things at
once: the rail's flag on (flags.require_rail), dry_run=False per call,
and an injected transport (flags.require_transport). Tests therefore
cannot reach the network by construction, and the shipped state (all flags
false, dry_run defaults true, no transports wired) is inert three ways.
GA4 goes further: send() only accepts a ValidatedPayload that only
validate() can construct — send-without-validate is a TypeError.
7. 48h Meta dedup window to spec
meta_capi.DEDUP_WINDOW_HOURS = 48 matches Meta's documented
event_id/event_name dedup horizon for pixel/CAPI pairs; the shared id shape
is mr-{tenant}-{source_event_id}. A unit test asserts the literal 48 so
the window cannot drift silently; boundary tests cover 47h (duplicate),
48h and 49h (new).
7a. Persistence is a third lock, not a code path (amended 2026-08-09)
Amends §6's scope. As first built, the only recompute surface was
POST /v1/rails/recompute:dry-run against an in-memory writer:
MeasurementRailsBQWriter existed and was tested but was instantiated
nowhere in production code, so no persisted row could ever be written and
causal_credit had no HTTP caller at all. The tables and *_latest views
were real; nothing could fill them.
The wiring adds POST /v1/rails/recompute and
POST /v1/rails/causal-credit:classify WITHOUT touching the attribution or
classification logic (same recompute(), same 7d-click/1d-view window, same
classify() with its deterministic-only guard) — the routes differ from the
dry-run route only in which writer they are handed.
Persistence borrows §6's shape rather than inventing a new one: a live
platform send needs flag + dry_run=False + injected transport, so a
persisted write needs MEASUREMENT_PERSIST=true + per-call "persist": true
+ a configured writer. Rejected alternative: making persistence the default
whenever GCP_PROJECT happens to be set — deployment configuration is not
consent to write, and it would make a misrouted scheduler job a data-writing
event.
A half-open configuration REFUSES (403 flag off, 503 no writer, 502 insert
failure) instead of falling back to the in-memory writer. A caller that asked
to persist and received "written": 1 from an in-memory writer would have
been told rows landed in BigQuery when they did not.
MEASUREMENT_PERSIST is deliberately NOT in the §6 outbound ladder and is
kept semantically distinct from MEASUREMENT_WRITEBACK: it governs
first-party storage in the tenant's own dataset and can never cause an
outbound platform send. It is reported in flag_states() (so /health and
/metrics show it) and its False default is pinned as source text by
test_flags.py, like the other defaults; the ambient sentinel now also
turns the suite red if the suite is invoked with it enabled.
7b. Drift streak state is durable, and says when it is not (2026-08-09)
The drift monitor held streaks in a process dict, so a Cloud Run instance
recycling on day 2 of a 3-day drift silently reset the count toward the
ALERT_CONSECUTIVE_DAYS = 3 threshold. An alert that never fires is
indistinguishable from "no drift" — a fabricated measurement by omission.
An injected observation store (bq.DriftObservationStore, the same
append-only injected-client pattern as MeasurementRailsBQWriter, plus the
read half needed to rehydrate) makes record_day append and hydrate()
rebuild the streak at construction. Re-recording a day appends; the read
keeps the newest row per (tenant_id, source, day) by recorded_at —
latest-wins, mirroring the *_latest views. No UPDATE path was added.
The store stays OPTIONAL because it needs operator-applied DDL
(RUNBOOK §1a) and credentials, so the honest posture is to report the
difference rather than imply durability: status() and /health carry
durable / drift_state_durable, false when the streak is memory-only.
A store append that fails RAISES (502 from the route) — a day the caller
believes was recorded but that no restart can recover is silent data loss.
write_alerts had an injected sink but no caller; the record route now
invokes it, defaulting to a structured Cloud Logging line. That default was
chosen over a webhook because it requires no credential and no HTTP client.
The constraint that must keep holding is a property of the CODE: no
measurement-rails module imports or constructs an HTTP client — enforced
by test_flags.py::TestNoHttpClientInRailsModules, which scans every
non-test module in the service directory.
Correction (recorded rather than quietly edited): an earlier revision of
this section stated the constraint as "requirements.txt contains no HTTP
client at all". That is false. requirements.txt names no HTTP library
directly, but the pinned google-cloud-bigquery==3.27.0 pulls requests
and urllib3 into the image transitively, so the image DOES ship an HTTP
stack. The behavioural property was and is true; the stated basis for it was
not. Dependency absence was never the real guarantee anyway — §6's
transport injection is: a send needs a flag on, dry_run=False, and an
explicitly injected transport, none of which an installed library provides.
Consequence: measurement_drift_observations is a new table whose DDL
currently lives in RUNBOOK §1a rather than
bigquery/schemas/measurement_rails_ddl.sql. Folding it in is a tracked
follow-up (RUNBOOK §8).
8. Bitemporal honesty is structural
house_attribution and causal_credit stamp available_to_model_at at
recompute/classification time, never business time, and RAISE
BitemporalViolation when it would precede observed_at (Article 6.1).
The causal ledger is append-only — the service has no UPDATE path
(Article 6.2).
9. Registry before exposure
design_registry requires mde up front, bounds holdout_pct to
[0.10, 0.20] for intent-driven activation designs (Cell 36 ships 10%
permanent holdout), enforces write-once (409 on re-register), and refuses
designs whose exposure already started — pre-registration that can be
back-dated is not pre-registration.
Consequences
- Merges to main build nothing, deploy nothing, schedule nothing for this
service; the operator owns every activation step (RUNBOOK order:
validate-only → dry-run → live per rail per engagement;
MEASUREMENT_WRITEBACKstrictly last, owner-approved). Persistence (§7a) is a separate, first-party-only decision the operator makes in RUNBOOK §5a, and it is off in the shipped deploy command. - Because no deploy workflow matches this service, the routes added in §7a are built and tested, not deployed — they exist in code and in the test suite, and are unreachable in production until an operator deploys (RUNBOOK §2/§8). No claim above "built, pre-benchmark" is available for any of it.
- Grafana alerting for drift is a stub
(
docs/measurement-rails/grafana-drift-alert.stub.yaml) until the operator wires the Prometheus datasource — placing it undermonitoring/would have entangled it with the live monitoring tree. - Numbers in
lag_profiles.yamlandaem_priority_schema.yamlare labeled operating defaults, replaced per-engagement with empirical fits; no claim stronger than "built, pre-benchmark" exists anywhere in this service's metadata, and none may be upgraded without benchmark artifacts (TRUTH.md Article 3).