Fleet Integrity — Multi-Merchant Auction Independence and Data Residency
Status: Active spec (v1.0, 2026-08-11) · Owner directive: Master v4.0 §3.5 "Fleet integrity",
review finding N4 (High — legal/architectural) · Lane: MIZOKI Signal for Shopify
Rule of precedence: this spec must be satisfied before pooled priors ship (P4 / B4).
Nothing in it is a performance claim; the platform ceiling stays built, pre-benchmark (TRUTH.md).
Sources: docs/product/SIGNAL_SHOPIFY_MASTER_v4.md §2.7, §3.5;
docs/product/SIGNAL_OVERVIEW_v5.md §2.6; owner decisions 5, 14, 15 (master §3.6).
Why this spec exists
As merchant count grows, MIZOKI merchants will meet in the same ad auctions and product categories. Two structural exposures follow: (1) pooled cross-merchant learning could leak one merchant's strategy to another, and (2) anything that even resembles coordinated bidding between merchants is auction-collusion legal exposure. Both are closed by construction, not by policy promise, and the immutable ledger is the proof surface.
(a) Pooled priors are DP-aggregated — merchant strategy is non-inferable
- Pooled category priors (T1 cold-start seeding; [Roadmap], P4/B4 per
docs/roadmap/SIGNAL_SHOPIFY_PHASE_BINDING.md) are computed only as aggregates across merchants under differential privacy; per-merchant contributions are bounded and noised such that no other merchant (or the platform's own serving path) can infer an individual merchant's bids, budgets, creative strategy, margins, or customer set from a prior. - Federated-learning transport with DP aggregation is the mechanism of record (master §2.7). Synthetic users never leave the training plane, never impersonate real people, never reach ad platforms.
- Priors are always labeled as priors on every surface that shows them (T1 receives "pooled priors (labeled as priors) — no merchant-level incrementality claims", master §1.3).
- Consent default for contributing to pooled priors is owner decision 5 (counsel) — open; no pooling ships before it is decided.
- [IN BUILD] No pooled-prior pipeline exists today. The DP parameters (ε budget, clipping bounds, minimum-cohort floor) must be specified and reviewed in the P4 build before any cross-merchant aggregate is computed, and a seeded inference-attack test must ship with the pipeline (both directions: the leak it must prevent, the aggregate it must still allow).
(b) No cross-merchant bid coordination — independence by construction, ledger-provable
- The platform never coordinates bids, budgets, pacing, or auction participation between
merchants. Each merchant's optimization loop runs independently: its own signals, its own
NCM objective (NCM-v1 metric contract), its own clipped-ReLU DEL authorization
(
docs/architecture/DEL_AUTHORIZATION_FUNCTION.md), its own covenant caps. - No optimization input may carry another merchant's contemporaneous auction state, bid, or budget. The only cross-merchant artifact permitted anywhere in the decision path is the DP-aggregated prior of (a) — historical, aggregated, non-attributable, and labeled.
- Proof surface: every proposed / authorized / denied action lands in the immutable action ledger with its reasoning path and DEL margin (master §2.7). Because inputs are journaled per decision, independence is audit-provable from the ledger: an auditor can verify that no decision for merchant A consumed merchant B's data. This is the artifact answering any auction-collusion inquiry — architecture, not assurance.
- [PARTIAL — test-pinned 2026-09-02, Wave 2 WS-9] The standing "independence audit"
query exists:
contracts/mizoki_contracts/independence_audit.py(audit_tenant/audit_fleet) walks a tenant'sdecision_proofs→ bounddecision_requests→validation_passports→action_authorizations/outcomes/approvals/learning_records→ resolvablesignal_refs, and returns counts plus a violations list naming collection, document and field. Read-only over the existingStore; an empty tenant isvacuous, never clean; a joined row with no tenant field isunattributed, never clean. Pinned in both directions (clean two-tenant ledger; seeded cross-tenant passport / bound request / signal ref) bytests/governance/test_fleet_integrity.py. Not yet done: the release-time job that RUNSaudit_fleeton every release (no CI step or schedule exists — it stays [IN BUILD]); opaquesig:…signal refs are reportedunresolved, not checked, until signal refs name ledger documents. - Bid / scoring path independence — what is test-pinned today
(
tests/governance/test_fleet_integrity.py): the growth-scheduler F4 → MMM prior path builds tenant A's posterior from A's clean cycles only, althoughDURABLE_STORE.list("f4_cycles")is an unfiltered collection scan — the invariant holds because of the per-cycletenant_idfilter on the mmm-propose path, and the suite pins that line by name; Cell 34 scoring reads are tenant-keyed at the SQL boundary (every identity / topic / horizonWHEREcarriestenant_id = @tenant_id), so its ONE shared transformer checkpoint (the weights env path read once insrc/cells/cell34/scoring_cell/config.py) is a shared MODEL, not a cross-tenant SIGNAL — a distinction, not a waiver. Two paths carry no tenant parameter on their inputs and are pinned only where safe: the lift-engine estimate requests (UpliftRequest,CATEPredictionRequest) name no tenant (its evidence output is tenant-stamped; the engine is imported by no governed service), and the boss-sideads_causal_controller.pyhas zero tenant references and is imported by no governed service (execution adapters default OFF). Both are open findings indocs/reports/WAVE2_PACKET_WS-9.mdwith proposed minimal fixes, not closed items.
(c) EU data residency — options memo for counsel (owner decision 14)
The platform runs in us-central1. EU-merchant onboarding requires a transfer-mechanism
determination before EU go-to-market (master §3.5; currently unscheduled). Options for
counsel, with engineering consequences:
| Option | Mechanism | Engineering consequence | Notes for counsel |
|---|---|---|---|
| 1. SCCs + supplementary measures | EU→US transfer under Standard Contractual Clauses; HMAC-tokenized identifiers (pseudonymized personal data, GDPR/CCPA locked decision) + Secret-Manager-held keys as supplementary measures | Lowest engineering cost: no new infrastructure; DPIA + transfer-impact assessment documents the pseudonymization boundary | Fastest path; carries Schrems-line legal-review burden per assessment cycle |
| 2. Regional processing (EU region) | EU merchant data processed in an EU region (e.g. europe-west*) end to end |
Highest engineering cost: regional duplicates of the ingest → BigQuery → graph → ledger path; per-region Cloud Run fleet, per-region secrets; cross-region control plane must carry zero personal data | Strongest residency posture; multiplies operating surface — sequencing decision vs. EU demand |
| 3. Hybrid (regional landing + tokenized processing) | Raw identifiers land and are tokenized in-region; only HMAC-tokenized, non-reversible records cross to us-central1 |
Moderate cost: a small EU tokenization/consent front end; the core stack stays single-region | Middle path — counsel must confirm tokenized records qualify as adequately safeguarded for transfer |
- The choice is owner decision 14 with counsel; engineering does not pre-empt it. No EU merchant is onboarded before the determination is recorded.
- Residency per tenant config — [IN BUILD — owner decision D-14 OPEN] (
docs/OPEN_ITEMS.mdD-14). The honest current contract, test-pinned bytests/governance/test_fleet_integrity.py(2026-09-02): NO residency field exists in any tenant config surface — not in the onboarding validator's*_KEYSsets (services/service-marketing-connectors/tenant_economics.py), the lane-config reader (contracts/mizoki_contracts/tenant_lane_config.py), or the onboarding page — so nothing can silently claim residency; and the deployment is single-regionus-central1by deploy-time constant (everydeploy-*.ymlservice region). The suite fails if a residency field appears while D-14 is still OPEN, or if D-14 closes while this label stands. GA4 Measurement Protocol EU-endpoint selection is by end-user country at send time, not a tenant residency setting. - Whatever the choice,
customers/redact/shop/redactwebhooks, the erasure cascade, and Cell 33's consent gate apply identically in every region (constitution II.6).
Enforcement hooks
- This spec gates the P4 exit ("pooled priors" cannot appear in any build plan without citing
it) — binding via
docs/roadmap/SIGNAL_SHOPIFY_PHASE_BINDING.md. - Merchant-facing copy about fleet learning must carry the (a) labeling rules. Discharged
2026-09-02 (Wave 2 WS-9): the first such copy shipped (marketing whitepaper r2.0 §10
"Fleet integrity" bullet; the Shopify offering §7 paragraph via the WS-9 packet), and the
content-QA rule shipped with it —
scripts/mizoki_canon.pyrule V14 (an unlabeled pooled / shared / federated prior described across merchants, or a coordination claim about bids, budgets, pacing or auctions between merchants — the negative framing and the "labeled as priors" framing stay legal), seeded in both directions intests/skills/test_canon_vocabulary.pyand enforced overdocs/marketing/*.md+docs/product/*.mdbyscripts/check_canon_docs.pyand over the skill corpus byscripts/skill_sync.py --audit.