SUPERSEDED by docs/product/SIGNAL_SHOPIFY_MASTER_v4.md
Historical fragment — committed under the standing rule in
README.md("if a fragment copy resurfaces, commit it here byte-for-byte with a supersession header prepended"). This is the "Product Definition v2.0 / v2.1" element of the recorded supersession chain, consolidated intoSIGNAL_SHOPIFY_MASTER_v4.0 FINAL(2026-08-11) underSIGNAL_OVERVIEW_v5.md(content v5.3). It resurfaced as an owner-supplied copy on 2026-08-21 and is preserved verbatim below — pre-master vintage claims are deliberately NOT corrected in place. Deltas a reader must take from current canon instead: (a) the intent graph is Firestore-backed — Neo4j was retired by owner decision 2026-08-09 (§3.1 below predates that); (b) holdout registration / causal credit is Cell 36intent-causal(Cell 35 =intent-graph; owner-directed correction 2026-08-19 recorded in the master's header — §§3.3/4.2 below carry the pre-correction numbering); (c) §10 questions 1–2 were DECIDED 2026-08-12 (master §3.6a: D1 = direct/unlisted distribution P1–P3 with App Store at P4; D2 = merchant-owned ad accounts only). Cite the master or the overview, never this fragment.
MIZOKI 3.5 × Shopify — Governed Media Acquisition for Mid‑Market and Emerging Merchants
Product Definition Document v2.1 (Enhanced, Integrated, Audit‑Corrected, Self‑Reviewed)
Status: Draft for owner review · Supersedes the "Democratizing Enterprise Decision Intelligence" draft Claim discipline: Every material claim in this document is tagged [Validated], [Illustrative], or [Roadmap] per the MIZOKI3 Content Constitution. Mixed-class claims are a defect. Vocabulary rules honored: MIZOKI3 is a nervous system, never a brain. Divisions are example deployments, not a product ceiling.
0. Executive Summary
The prior draft correctly identified the strategic thesis: independent Shopify merchants lose money to self-grading ad-platform attribution and revenue-based (rather than profit-based) optimization, and MIZOKI 3.5's governed SRPVDAL loop can close that gap. This v2.0 document keeps the thesis and fixes the four structural problems that would have prevented it from becoming a shippable product:
- Claim contamination. The draft presented research findings from third parties (Airbnb's synthetic-query KL-divergence numbers, Amazon COSMO architecture) in ways that read as MIZOKI results. v2.0 re-classes every claim and separates "research we build on" from "results we have."
- Missing product boundaries. No target-customer definition, no pricing logic, no minimum-viable data requirements. v2.0 defines three merchant tiers with honest capability floors — including where statistical incrementality testing is not yet possible and what the product does instead.
- Incomplete mechanics. The autonomy ladder table, the Net Contribution Margin formula, and the incrementality method list were truncated or absent. v2.0 specifies them fully, aligned to the shipped cells 33–36 (
intent-signal-ingest,intent-scoring-api,intent-graph,intent-causal) and the eightintent_*MCP tools. - No competitive or ecosystem reality. The draft ignored that Shopify itself (Audiences v2.4, Benchmarks), Triple Whale (Moby AI + Lighthouse), Northbeam, Sellforte, Rockerbox, and Google's open-source Meridian/GeoX all occupy adjacent ground. v2.0 positions MIZOKI in that field and identifies the one lane no incumbent holds: governed autonomous execution optimized to net contribution margin, gated by always-on causal measurement.
The product this document defines: MIZOKI Signal for Shopify — a Shopify app + backend deployment of the Signal division that ingests store telemetry into the canonical event envelope, scores intent via the live LII cells, measures true lift via a tiered incrementality stack, and executes budget/bid/creative actions across Meta, Google, TikTok, and Klaviyo under the Decision Control Plane, with spend authority earned level-by-level.
1. Audit of the Prior Draft — Issues Resolved in This Version
| # | Issue found | Class | Resolution in v2.0 |
|---|---|---|---|
| 1 | KL-divergence figures (4.95 → 0.66, attribute divergence 0.04) are from the Airbnb LLM synthetic-data paper (arXiv 2605.21812), presented as if they were Mizoki measurements | Integrity | Re-attributed as research foundation [Validated as third-party research]; MIZOKI's own synthetic-seeding performance is [Roadmap] until we produce our own eval artifact (see §6.4 — same baseline-then-claim protocol as the GraphRAG harness) |
| 2 | "Quokka Swarm Optimization (QSO)" as a validation method | Credibility | Removed. Validation panel specified with standard, defensible methods: uncertainty/confidence classifiers, distributional-divergence gates, and human-auditable rejection logs. Exotic algorithm names without citations violate the Content Constitution's tone rules |
| 3 | Autonomy ladder announced as "six delegation levels" but the table was missing | Completeness | Fully specified in §5.2 (L0–L5) with spend caps, calibration gates, and rollback rules |
| 4 | Net Contribution Margin formula truncated ("()") with orphaned "Where:" | Completeness | Full formula and variable definitions in §4.2 |
| 5 | Reference list duplicates the same 4 URLs across footnotes [1]–[7]; blanket citation of every paragraph | Integrity | v2.0 cites per-claim; illustrative numbers marked as such |
| 6 | Return-rate margin math ("$45 gross → $-2 to $16 net") stated as fact with no source | Integrity | Reclassed [Illustrative] with the mechanism explained; real per-merchant numbers come from the merchant's own COGS/3PL feed at onboarding |
| 7 | "15-minute execution cycles" asserted without tying to platform learning-phase constraints | Technical | §4.4 reconciles execution cadence with Meta/Google learning-phase and budget-change dampening rules — high-frequency sensing, dampened acting |
| 8 | No consent, privacy, or platform-ToS treatment for autonomous spend on merchant accounts | Risk | §7 adds the governance/compliance layer: GDPR/CCPA posture (HMAC-tokenized identifiers = pseudonymized personal data, per locked platform decision), Shopify protected-customer-data requirements, ad-platform automation ToS, and spend-liability framing |
| 9 | No target market definition despite "mid-market to lower market" being the stated goal | Product | §2 defines tiers with data floors and honest capability limits |
| 10 | "Superior to traditional human agency management" stated flatly | Tone | Softened to measurable comparison criteria; superiority claims only against defined metrics with artifacts (Content Constitution: Validated/Illustrative/Roadmap never mixed) |
2. Target Market and Honest Capability Tiers
The stated target is mid-market to lower-market Shopify merchants, specifically for media acquisition. Statistical power is the binding constraint — a merchant doing 40 orders/month cannot support geo-holdout experiments, and pretending otherwise is how measurement vendors lose trust. The product therefore tiers by data volume, not by plan price alone.
| Tier | Profile [Illustrative bands] | Monthly ad spend | Orders/mo | What MIZOKI can honestly deliver |
|---|---|---|---|---|
| T1 — Emerging | New store or new SKU line | <$5K | <300 | Profit-aware value feeds (CAPI/conversion value rules), synthetic cold-start seeding, guardrail governance (L0–L1 autonomy), pooled category benchmarks. No merchant-level incrementality claims — directional intent scoring only |
| T2 — Growth | Established DTC | $5K–$50K | 300–3,000 | Everything in T1 + ghost-bid and intent-holdout experiments (user-level), rotating always-on holdout (5–15%), NCM-optimized budget reallocation, L2–L3 autonomy |
| T3 — Mid-market | Multi-channel, 3PL-integrated | $50K–$500K+ | 3,000+ | Everything in T2 + geo-lift experiments (Meridian GeoX-class methodology), mini-MMM calibrated by lift tests, cross-channel reallocation, L4–L5 autonomy |
Why this matters for research grounding [Validated as third-party research]: the market has converged on a four-rung "attribution confidence ladder" — (1) platform self-reported ROAS, (2) third-party pixel attribution, (3) first-party pixel + CAPI + Shopify reconciliation, (4) incrementality testing + blended MER. Most tools sell Rung 2. MIZOKI's product promise is to move every merchant to the highest rung their data volume supports, and to say which rung they're on.
3. Architecture: Shopify-Native Integration into the Canonical Event Envelope
3.1 Ingestion (SENSE)
- Shopify App (embedded) + OAuth scopes: orders, products, inventory, customers (protected-data tier), fulfillment, price rules. Webhooks for
orders/*,refunds/*,inventory_levels/*,fulfillments/*; bulk operations for backfill. - Web Pixel Extension: session telemetry (product views, add-to-cart, checkout steps) →
intent-signal-ingest(Cell 33) through the consent gate. Consent state travels with every event. - Cost-side feeds: landed COGS (per-variant), 3PL pick/pack/dimensional surcharges, payment-gateway fees, return-processing costs. Sources: Shopify cost-per-item field, ERP/3PL connectors, or a guided onboarding worksheet for T1 merchants who have none of these systems.
- Ad-platform feeds: Meta, Google, TikTok spend/impression/click data; Klaviyo flows for owned-channel context.
- All streams normalize into the Canonical Event Envelope and land in BigQuery
unified+ the Neo4j intent graph — the same envelope the rest of the platform speaks. No Shopify-special data model. [Validated — platform architecture as shipped]
3.2 Intent layer (REASON)
- Session trajectories map onto the Intention Graph via the IGC-RC framing (intention generation → conceptualization → relation classification), with asynchronous / synchronous / causal edge types. This is the enterprise pattern established by Amazon COSMO and intention-KG research (arXiv 2412.11500) [Validated as third-party research]; MIZOKI's implementation is the live cells 33–36 [Validated — deployed and IAM-locked per July 31 measurement].
- Scoring surface:
intent_score_get(stage, next interests, confidence, explanation path),intent_cohort_query(in-market lists),intent_transitions_recent. - Differentiator vs. keyword/CPV catalogs: bids are placed against predicted downstream purchase goals before demand inflates auction prices — but every LII-driven activation requires a registered holdout first (
uplift_export_cohortis DCP-gated). Observe-only is the default. [Validated — platform governance rule]
3.3 Measurement layer (VALIDATE — the credibility core)
Tiered by merchant volume (§2):
- Ghost bidding — log the bid decision, randomly withhold execution for a control slice, compare outcomes. Cheapest experiment; works at T2 volumes. Registered via the holdout API (Cell 35, Incrementality & Causal Credit).
- Intent-cohort holdouts — suppress activation for a randomized slice of an in-market cohort; measures caused-vs-anticipated conversions (
intent_incrementality_report). - Geo experiments — Meridian GeoX-class methodology (time-based regression + stratified market matching, multi-cell against a shared control) [Validated as third-party methodology — Google has open-sourced it]. T3 only; requires geographic order density. Publisher-agnostic, which matters because platform-run "conversion lift" is still the platform grading itself.
- Mini-MMM calibration — a lightweight Bayesian channel-level model (Meridian-style adstock + saturation) calibrated by the lift tests above, for T3 budget planning. MMM without experiment calibration is correlation dressed up as insight; the product never ships an uncalibrated MMM number. [Roadmap]
- Pooled priors for T1 — small merchants inherit category-level priors from the (consented, aggregated) fleet, clearly labeled as priors, not measurements. [Roadmap]
Causal-inference machinery: X-Learner / DR-Learner uplift with DoWhy refutation (Cells 26–27), consistent with the platform's existing causal stack. [Validated — platform architecture]
3.4 Execution layer (ACT)
- Meta: CAPI server-side events with EMQ-maximizing hashed parameters (email, phone, external_id, fbp/fbc), shared
event_iddedup with the pixel; campaign objective "Maximize value of conversions" once eligibility clears; conversion value = net contribution, not revenue. Sending raw revenue trains the platform to find discount-hunters and refunders; sending margin/pLTV trains it to find profitable, high-retention customers. [Validated as third-party best practice; MIZOKI automation of it is the product] - Google: Enhanced Conversions + Conversion Value Rules carrying NCM adjustments; new-customer-acquisition goal where LTV economics justify it.
- Klaviyo: owned-channel suppression/activation synchronized with paid — never pay to reach someone the flow already converts (this is the paid-side twin of Shopify Audiences' Existing Customers Plus exclusion logic).
- Shopify Audiences interop (T3/Plus): MIZOKI complements Audiences rather than competing with it — Audiences supplies platform-side match lists; MIZOKI supplies the value signal, the profit gate, and the causal verdict on whether the Audiences list is actually incremental (Shopify itself recommends head-to-head experiments for exactly this).
4. Profit-Aware Bidding: From Reported ROAS to Net Contribution Yield
4.1 Why revenue optimization fails [Validated as mechanism; numbers Illustrative]
Platforms attribute conversions to their own placements regardless of causation, and human managers optimize GMV/AOV. In high-return categories (apparel: 25–40% online return rates), returns processing, freight, restocking, and markdowns can take a healthy gross margin per order to near zero or negative. A campaign can be "3.5× ROAS" and cash-flow negative. Merchant-specific numbers are computed at onboarding from their own COGS/3PL feeds — the product never asserts a generic figure.
4.2 Net Contribution Margin — full specification
Per order i, per cohort C:
NCM(C) = Σᵢ∈C [ Rᵢ − COGSᵢ − Fᵢ − Sᵢ − Pᵢ − E[RLᵢ] ] − AdSpend(C)
Where: - Rᵢ — net revenue after discounts and gift-card adjustments - COGSᵢ — landed component cost of goods (bundle-decomposed to component level — bundle/kit profit distortion is a known Shopify analytics failure mode) - Fᵢ — fulfillment: pick/pack + dimensional-weight shipping - Sᵢ — platform + payment processing fees - Pᵢ — promotional cost allocation (site-wide discounts pro-rated) - E[RLᵢ] — expected reverse-logistics cost: predicted return probability (SKU × cohort × season) × (return freight + processing labor + restock/markdown loss). Modeled, then trued-up when the actual return window closes - AdSpend(C) — spend attributed to cohort C under the causal credit model (Cell 35), not last-touch
Secondary objective for T2+: incremental NCM per marginal dollar (miNCM) — the derivative that decides where the next dollar goes. This is the metric almost no incumbent computes at campaign level (Triple Whale's own gap is the absence of marginal incremental ROAS/saturation curves). [Validated as competitive gap per public documentation]
4.3 Value feedback loop
E[NCM] per conversion → CAPI / Conversion Value Rules → platform bidders optimize toward margin-positive cohorts. The platforms' own ML becomes MIZOKI's execution substrate; MIZOKI owns the objective function and the causal audit.
4.4 Execution cadence — corrected
- Sense/score: continuous (sub-second canonical streaming, as platform SLOs already target).
- Act: dampened. Budget/bid moves respect platform learning-phase mechanics — large or frequent budget changes reset learning and destroy the very performance being optimized. Default action windows: creative rotation and SKU pause/resume near-real-time; budget reallocation evaluated up to 4× daily but with cumulative change capped at ±20% per campaign per day (matching the L3 authority limit in §5.2 — one number, stated once, enforced everywhere); structural changes weekly. The "15-minute loop" survives as the decision evaluation cadence, not a guaranteed action cadence. Every act passes SRPVDAL Validate → Decide first.
4.5 The black-box campaign reality (Advantage+ / Performance Max) — added in v2.1
The mid/lower market in 2026 predominantly runs Meta Advantage+ Shopping and Google Performance Max — campaign types that deliberately remove granular audience and placement control. A product spec that assumes hand-built campaign structures is optimizing a shrinking surface. In the black-box era, the merchant's remaining levers are exactly four: (1) the value signal (what "a conversion is worth" — our NCM feed), (2) exclusions and suppressions (existing customers, high-return cohorts), (3) creative supply, and (4) budget and guardrails. MIZOKI Signal is therefore designed lever-native: it does not fight the platform bidder for control it can't have; it owns the objective function the bidder trains on, the exclusion lists, the creative rotation, and the covenant caps. Manual-structure micro-management remains available for T3 merchants who run it, but the default product story is "we steer the black box by controlling its inputs and auditing its outputs causally." [Validated as ecosystem direction; implementation Roadmap]
4.6 Creative intelligence layer — added in v2.1
Once value signals converge, creative becomes the largest remaining performance variable, and it was nearly absent from the prior spec. Scope for MIZOKI Signal: - Fatigue detection: frequency-adjusted response decay per creative per cohort from the canonical stream; rotate before CPM/CVR degradation, not after (this operationalizes the one-line mention in the draft). - Cohort-creative fit: intent-graph stages mapped to message archetypes (problem-aware vs. product-aware vs. offer-ready) — the intent layer's first direct creative application. - Structured testing: creative variants enter as SRPVDAL plans with holdout-aware evaluation, so "winning creative" means causally-verified NCM lift, not click-through vanity. - Generation is out of scope for GA [Roadmap decision]: MIZOKI scores, schedules, and retires creative; it does not generate it at launch. This keeps the app out of brand-safety liability and keeps the moat on governance + measurement. Revisit post-Phase 3.
4.7 Shopify-native demand channels: Shop Campaigns and feed quality — added in v2.1
- Shop Campaigns interop [Roadmap]: Shopify's own CAC-target acquisition channel (merchant sets a target cost, pays on results) is a natural fit for lower-market merchants and an honest benchmark competitor for our own acquisition claims. MIZOKI treats it as one more channel in the reallocation set, with its results subject to the same causal audit as Meta/Google — including the question of whether Shop-attributed orders are incremental.
- Catalog/feed enrichment: PMax and Advantage+ delivery quality is materially a function of product-feed quality (titles, attributes, categorization). The intent graph + the platform's ontology tooling can enrich feed attributes from observed search-and-session language — a low-risk, high-visibility win for T1 merchants that needs no spend authority at all (works at L1).
4.8 Working-capital-aware pacing — added in v2.1
For sub-$50K/mo merchants, the binding constraint is often cash and inventory cover, not auction efficiency. MIZOKI Signal reads inventory positions and sell-through velocity (already in the canonical stream) and paces acquisition spend against weeks-of-cover and a merchant-set cash floor: never spend to sell inventory that isn't there; throttle before stockouts strand learning-phase investment; re-accelerate on replenishment. This connects the Media Acquisition cell to the Predictive Financial domain cell — the first concrete cross-domain hook in the Shopify product, and a differentiator no attribution vendor offers because none of them carry a financial domain model. [Roadmap; architecture Validated]
5. Governance: SRPVDAL and the Autonomy Ladder (fully specified)
5.1 The seven-phase loop applied to media
Sense (canonical events) → Reason (graph + intent + causal memory) → Plan (candidate reallocations/rotations/pauses) → Validate (counterfactual simulation against margin guardrails, inventory positions, promo calendar, active experiments — validation labs are non-bypassable per locked platform decision) → Decide (DEL score vs. policy threshold; below threshold → human approval queue with reasoning path) → Act (idempotent, signed API calls; every action journaled) → Learn (prediction-vs-outcome deltas update edge weights and return-probability models).
5.2 Autonomy Ladder — six delegation levels
| Level | Name | Authority | Promotion gate | Demotion trigger |
|---|---|---|---|---|
| L0 | Observe | Read-only; recommendations with reasoning paths | Default at install | — |
| L1 | Signal | Push value signals (CAPI/CVR) and audience syncs; no spend changes | 14 days clean data reconciliation (Shopify orders vs. tracked events within tolerance) | Data-quality breach |
| L2 | Housekeep | Pause low-inventory/high-return SKU ads; rotate creative; apply exclusions | 1 full purchase cycle at L1; zero guardrail violations | Any unexplained NCM regression |
| L3 | Reallocate | Move budget between campaigns within a platform, ±20%/day, under a hard monthly cap set by the merchant | Calibrated forecast accuracy over 2 purchase cycles; ≥1 completed holdout | Forecast error beyond band; guardrail breach |
| L4 | Cross-channel | Reallocate across platforms; adjust bid strategies/targets | ≥2 completed incrementality experiments; DEL threshold history clean | Experiment invalidation; merchant override |
| L5 | Autonomous yield | Full media-acquisition authority within covenant caps (max daily spend, min cash-margin floor, category exclusions) | T3 data volume + sustained miNCM accuracy + explicit merchant covenant signature | Any covenant proximity event auto-drops to L3 |
Every level: instant one-tap kill switch, immutable action journal, and weekly plain-language digest of what was done and why. Spend authority is earned through calibration, never granted by plan tier. [Validated — matches Decision Control Plane design]
6. Solving Cold Start Honestly
6.1 The problem
New stores and new SKUs lack the interaction history that collaborative-filtering and platform learning phases need; the result is multi-week exploration burn at elevated CAC.
6.2 The research foundation [Validated as third-party research]
LLM-based synthetic interaction generation with contrastive seed guidance materially closes the realism gap: Airbnb's published work reports unguided generation at KL divergence ≈4.95 vs. real query-length distributions, reduced to ≈0.66 with seeding (attribute divergence ≈0.04). Two-Tower retrieval architectures warm-started on filtered synthetic data are the standard pattern.
6.3 The MIZOKI implementation [Roadmap until eval artifact exists]
- Contrastive item pairing + empirical seeds from (a) the merchant's earliest real sessions, (b) consented pooled category priors.
- Validation panel: distributional-divergence gates against real-traffic references, uncertainty classifiers, and an auditable rejection log. (No exotic unnamed-provenance algorithms — credibility is the product.)
- Output warm-starts intent-graph priors and CAPI seed audiences on Day 1, shortening — not skipping — platform learning phases.
6.4 Claim protocol
Per the platform's baseline-then-claim rule (same discipline as the GraphRAG eval harness): no cold-start performance claim ships until a frozen gold-set baseline and post-change artifact exist. Marketing may describe the mechanism [Illustrative] but not assert lift numbers until then.
7. Compliance, Consent, and Platform Risk (absent from prior draft)
- Privacy posture. HMAC-tokenized identifiers are treated as pseudonymized personal data in GDPR/CCPA scope (locked platform decision). Cell 33's consent gate applies to pixel telemetry; GDPR subject-access/erasure paths already exist (cells 33–36 erasure cascade). Shopify protected-customer-data requirements: request minimum scopes, document purpose, honor redaction webhooks (
customers/redact,shop/redact). - Ad-platform ToS. Automated management through official APIs (Meta Marketing API, Google Ads API) with merchant-authorized tokens — never credential automation of Ads Manager UIs. Value adjustments via sanctioned mechanisms (CAPI, Conversion Value Rules) only.
- Spend liability. Autonomy covenant is a signed instrument: caps, floors, exclusions, and the demotion ladder are contractual, not just UI settings. Counsel review required before L4/L5 GA. This is the commercial version of the DEL demo story (ACT-991: proposed action re-routed to the smaller option when the DEL score missed threshold) — the pitch is the governance.
- Synthetic-data hygiene. Synthetic interactions never impersonate real identifiable users, never leave the training plane, and are never sent to ad platforms as audience members.
- Benchmark substantiation. Any public performance claim (e.g., against the 40% CAC-reduction platform goal) requires named-metric artifacts. The FTC substantiation bar applies to "up to X%" claims.
8. Competitive Position
| Capability | Shopify Audiences v2.4 | Triple Whale (Moby + Lighthouse) | Northbeam / Rockerbox | Sellforte | MIZOKI Signal for Shopify |
|---|---|---|---|---|---|
| Audience match lists | ✅ core | — | — | — | Interop, not replicate |
| Attribution dashboards | Benchmarks only | ✅ core | ✅ core | via MMM | Included, not the point |
| Incrementality testing | merchant-run DIY | ✅ geo (Lighthouse) | tiered/enterprise | ✅ campaign-level iROAS | ✅ tiered: ghost-bid → cohort holdout → geo → MMM-calibrated |
| Marginal iROAS / saturation | — | ✗ (documented gap) | partial | ✅ | ✅ miNCM as first-class objective [Roadmap for GA] |
| Profit (net contribution) objective | — | profit dashboards | — | revenue-centric | ✅ NCM with modeled reverse logistics, bundle-decomposed COGS |
| Autonomous execution | — | ✅ push-to-platform, configurable autonomy | — | insight-only | ✅ governed — DEL-gated, covenant-capped, earned autonomy ladder |
| Intent graph / pre-demand bidding | — | — | — | — | ✅ cells 33–36 live |
| Governance artifacts (journals, reasoning paths, kill switch) | — | partial | — | — | ✅ core identity |
The defensible lane: competitors have either measurement without execution (Sellforte, Northbeam) or execution without causal grounding (Triple Whale's recommendations aren't incrementality-based, per its own public documentation). Nobody combines profit-objective + always-on causal gating + governed autonomy — and governance is the hardest part to retrofit because it's an architecture, not a feature. That is the SRPVDAL moat.
Pricing posture [Roadmap, for owner decision]: T1 self-serve (~$99–$299/mo range of the category), T2 percentage-of-spend or flat tier undercutting agency retainers, T3 platform pricing. The category comparison set prices $199 add-ons for CAPI alone and ~$2K/mo entry for enterprise measurement — a governed execution engine sits above measurement-tool pricing and below agency cost.
9. Implementation Roadmap
Phase 1 — Foundation (T1 shippable). Shopify app + OAuth + webhooks + web pixel extension; canonical envelope mapping; COGS onboarding worksheet; NCM computation; CAPI/CVR value feeds; L0–L1 autonomy; weekly digest. Exit criterion: order-event reconciliation within tolerance across 10 design-partner stores.
Phase 2 — Measurement (T2 shippable). Holdout registration UX on the existing Cell 35 API; ghost-bid + intent-cohort experiments; causal credit reports in merchant-readable language; L2–L3 autonomy with the calibration gates of §5.2. Exit criterion: first third-party-auditable incrementality artifact per design partner.
Phase 3 — Yield (T3 shippable). Geo experiments (GeoX-class), mini-MMM calibrated by lift tests, cross-channel reallocation, miNCM optimization, L4–L5 covenants with counsel-approved instrument. Exit criterion: sustained miNCM forecast accuracy across 2+ purchase cycles at ≥3 T3 merchants.
Phase 4 — Scale. Cold-start synthetic seeding GA (post eval-artifact), pooled category priors, Shopify App Store listing + Built-for-Shopify badge, Audiences interop for Plus merchants.
Product KPIs (added in v2.1 — the spec previously defined merchant metrics but none for the product itself): - Activation: % of installs reaching L1 (clean reconciliation) within 14 days - Time-to-first-artifact: median days from install to first completed incrementality experiment (T2+) - Truth delta: median gap between platform-reported ROAS and MIZOKI-measured incremental NCM per merchant (the number the "Profit Truth Audit" wedge sells) - Autonomy trust: % of DEL-cleared actions executed without merchant override; demotion events per 100 merchant-months - Retention proxy: % of merchants advancing at least one autonomy level per quarter
Coordination note: this document should be handed to the designated coordinator session ("MIZ OKI 3.5 Shopify media buying," owner of the Shopify/media/net-yield/marketing-docs/JourneyEvent lanes) as the product source-of-truth, with features claimed in the coordination ledger before build — the PR #644 duplicate-work lesson applies.
10. Questions You Didn't Ask (and Should Decide)
- Distribution: App Store vs. direct sales? App Store review constrains scopes and adds Shopify's tax but gives the mid/lower market discovery. Recommendation: App Store for T1/T2, direct for T3. Requires deciding now because protected-customer-data approval shapes the pixel design.
- Who holds the ad accounts? Merchant-owned accounts with authorized tokens (recommended — clean liability) vs. MIZOKI-managed accounts (agency model, better margins, heavier compliance). This decision gates the L5 covenant language.
- What is the free wedge? A read-only "Profit Truth Audit" (connect store + ad accounts → 7-day report showing reported-ROAS vs. NCM vs. estimated incrementality) is a near-zero-cost L0 product that converts on its own findings. Strong candidate for the T1 top-of-funnel.
- Return-window accounting. NCM is provisional until return windows close (30–60 days in apparel). Decide the true-up policy: retroactive value adjustments to platforms (supported via CAPI) vs. forward-only correction.
- Fleet-learning consent. Pooled category priors (T1's cold-start lifeline) require an explicit data-sharing clause. Opt-in default or opt-out default is a legal + trust decision — counsel input needed.
- Klaviyo/email cannibalization measurement. Paid vs. owned-channel incrementality is the most common hidden money leak; deciding whether email suppression experiments are in Phase 2 or 3 changes the Klaviyo integration depth.
- What happens when MIZOKI recommends spending less? The honest system will sometimes conclude a merchant's incremental yield is negative and the right action is cutting spend. Pricing on % of spend punishes honesty; flat/performance pricing aligns it. This interacts with #3 and the whole GTM.
- Naming. The draft alternates "Mizoki 3.5" / "MIZOKI3"; the site says MIZOKI3 with Signal as a division. Recommend the commercial name MIZOKI Signal for Shopify and reserve version numbers for engineering.
Appendix A — Research Foundation (properly attributed)
- Airbnb, "Bridging the Cold-Start Gap: LLM-Powered Synthetic Data Generation for Natural Language Search" (arXiv 2605.21812) — synthetic-query realism via contrastive seeding.
- "Intention Knowledge Graph Construction for User Intention Relation Modeling" (arXiv 2412.11500) — asynchronous/synchronous/causal intention edges.
- Amazon Science, COSMO — common-sense knowledge generation from behavioral logs.
- Google Meridian + Meridian GeoX — open-source Bayesian MMM and publisher-agnostic geo-incrementality (TBR + stratified sampling, multi-cell shared control).
- Shopify Audiences v2.4 — Retargeting Boost, Existing Customers Plus exclusion, Benchmarks; Shopify's own guidance to run head-to-head incrementality experiments.
- Public competitive documentation: Triple Whale (Moby autonomy spectrum; Lighthouse geo testing; documented absence of campaign-level miROAS), Sellforte (campaign-level iROAS), Northbeam/Rockerbox (enterprise measurement posture), category pricing signals.
- Platform-internal [Validated]: SRPVDAL seven-phase loop, Canonical Event Envelope, Decision Control Plane + DEL, cells 26–27 (X-Learner/DR-Learner/DoWhy), cells 33–36 live (LII ingest/scoring/graph/causal), holdout-before-activation rule, non-bypassable validation labs, GDPR erasure cascade.