LII Anticipatory Intent — Phase 0 Discovery Report

Date: 2026-08-18 Phase: P0 (read-only discovery, no code) Claim: Latent Intent Inference integration plan audit — cells 33–36 existing LII, lift-engine causal modules, schema gaps, cell allocation, build gates Plan of record: docs/INTENT_API_PLATFORM_BLUEPRINT.md (the existing blueprint, v1.0, 2026-07-30) External plan: Boris's "How 'They Read My Mind'" anticipatory-intent plan (LII_ANTICIPATORY_INTENT_PLAN_v1.md) — NOT FOUND IN REPO; must be sourced before build work begins.


0. Plan Document Status

The external plan document (LII_ANTICIPATORY_INTENT_PLAN_v1.md) was not found at outputs/, docs/, or anywhere in the repository. A search across the Google Drive working directories also returned no match. The plan must be provided by the owner before any build-phase work can proceed.

This discovery report audits the existing LII infrastructure against the plan's known scope (new cells for latent intent, identity/household graph, incrementality/causal; behavioral micro-signals; interest graph; caused-vs-anticipated credit), as described in the tasking context.


1. Cell Registry — Ground Truth

Source: docs/architecture/CELL_REGISTRY.md v1.1 (the number authority, owner rulings 2026-08-11/2026-08-16)

Existing LII cells (33–36)

Cell Service SRPVDAL band Status
33 intent-signal-ingest SENSE deployed, IAM-locked, dispatch-only; live in Shadow since 2026-07-31
34 intent-scoring-api REASON/serving deployed, IAM-locked, --min-instances=1; observe-only; model quality gate OPEN
35 intent-graph REASON deployed, IAM-locked; durable store phase 1 merged; DSAR live-verified 2026-08-10
36 intent-causal VALIDATE/DECIDE live-verified 2026-08-11, rev intent-causal-00012-j27 serving

Cell 37 (market-signal-ingest) is the Data Injector & External Intelligence Gateway — partially operational, external-intelligence scope largely [IN BUILD]. Cells 38–39 are CRE Prospecting, deployed-CI.

Proposed assignments for the LII plan's new cells

The plan proposes three new logical components: (a) a latent-intent cell, (b) an identity & household graph cell, and (c) an incrementality/causal cell. However, much of this capability already exists in Cells 33–36 and the lift-engine service. Before allocating new cell numbers, the owner must decide whether the plan's components are extensions of existing cells or new cells.

If new cells are needed:

Cell Proposed name Purpose Notes
40 lii-behavioral-signals Behavioral micro-signal ingestion (dwell, scroll, partial-watch, hover — the signals Cell 33 does not yet accept) Could also be an extension of Cell 33's payload contract
41 lii-identity-graph Identity resolution, household stitching, cross-device linking Cell 35 already does intent graph; this may be an extension
42 lii-interest-embeddings Two-tower embeddings, ANN retrieval, interest co-occurrence graph Currently described as Cell 34 Phase B; may warrant its own cell

Owner decision required (Article VI): whether to allocate new cells (40–42) vs. extend existing cells (33–36). The build-debt approach of extending existing cells is lower-risk and follows the owner's ruling on CRE (extend cells inline rather than creating a privacy cell — ADR-LII-001 decision (a)).


2. Existing LII Code Inventory

Grand total: ~43,749 lines across Cells 33–37 + shared package + SDK + MCP server, with 1,023 test functions. All cells except Cell 37 use exact dependency pins.

2.1 Cell 33 — intent-signal-ingest (SENSE)

Location: src/cells/cell33/ Key modules: - ingest_cell/orchestrator.py — consent gate (fail-closed), audio prohibition, signal taxonomy validation, envelope construction, BigQuery persistence - ingest_cell/main.py — FastAPI routes: POST /v1/signals, GET /health, GET /readyz, GET /metrics - schema/intent_bigquery_ddl.sql — DDL for mizoki_intent tables including intent_signals, intent_outcomes, intent_consent_denials, intent_taxonomy_dim, intent_erasure_audit

Governance enforcement: audio prohibition (3 layers), consent fail-closed (region-aware, EU/EEA/UK deny-by-default), sensitive-topic deny-list (boot refusal), payload contract closed (ALLOWED_PAYLOAD_KEYS = {"signal_type", "topic_id", "strength", "identity_kind", "consent"})

Lines: 6,894 (source ~3,500, tests ~3,400) | Tests: 192 functions | Dependencies: all exact pins

Status: PRODUCTION-READY (deployed, IAM-locked, Shadow since 2026-07-31)

2.2 Cell 34 — intent-scoring-api (REASON)

Location: src/cells/cell34/ Key modules: - scoring_cell/orchestrator.py — scoring pipeline, cohort queries, transition detection, explain proxy, DSAR cascade coordinator - scoring_cell/scoring.py — compute_features() (9 fv2 features), heuristic_score() (fallback), FEATURE_VERSION = "fv2" - scoring_cell/main.py — FastAPI routes: GET /v1/intent/score/{identity_id}, POST /v1/intent/scores:batch, GET /v1/intent/cohort, GET /v1/intent/transitions, GET /v1/intent/explain/..., GET /v1/intent/taxonomy, POST /v1/intent/subject/{identity_id}:erase - backtest/ — offline evaluation harness with fv2 model, oracle ceiling analysis, bootstrap CIs - model/bqml.sql — production BOOSTED_TREE_CLASSIFIER definition

Model quality (from docs/reports/LII_BACKTEST_MODEL_QUALITY_2026-08-16.md): - Baseline (heuristic): AUC 0.6726 mean / Brier 0.2224 - fv2 model: AUC 0.6967 mean / Brier 0.1845 (crosses Brier ≤ 0.20 gate on 8/8 seeds) - Oracle ceiling: AUC 0.7277 (the 0.72 gate is unreachable on synthetic data) - Binding constraint: no forward labels — blocked on real merchant data (Shopify webhooks, GA4, Klaviyo)

Lines: 11,959 (source ~4,100, tests ~4,200, backtest ~2,740) | Tests: 275 functions | Dependencies: all exact pins

Status: PRODUCTION-READY (deployed, IAM-locked, observe-only; Phase A ranking gate FAILED and OPEN)

2.3 Cell 35 — intent-graph (REASON)

Location: src/cells/cell35/ Key modules: - graph_cell/graph_store.py — 1,759 lines; in-memory graph backend with frozen namespace - graph_cell/orchestrator.py — graph operations, explain paths, DSAR erasure leg - graph_cell/sync.py — score-to-graph synchronization - graph_cell/durable.py — optional Firestore write-through durability journal

Graph model (implemented):

Node type Key Purpose
IntentIdentity (tenant_id, identity_id) Deterministic identities only
IntentTopic (topic_id) Taxonomy topics (shared across tenants)
SignalAggregate (agg_id) Per-identity signal-type aggregates
IntentExplanation (explanation_id) Score explanation paths
Edge type From → To Purpose
HAS_INTENT Identity → Topic Score, horizon, in_market flag
HAS_EXPLANATION Identity → Explanation Explain proxy
CITES Explanation → SignalAggregate Contribution ranking
EXHIBITED Identity → SignalAggregate Weight, last_seen

What is NOT implemented: - PRECEDES co-occurrence edges (described in OFFERING_MAP F.4 and blueprint §3 Cell 35) - Household/cross-device linking nodes - Interest embeddings or vector indices - Two-tower candidate generation paths

Backend: in-memory; Neo4j permanently retired (2026-08-09 owner decision, NXDOMAIN by choice). Firestore durability journal available but off by default.

Lines: 11,127 (source ~4,990, tests ~5,740) | Tests: 321 functions | Dependencies: all exact pins

Status: PRODUCTION-READY (deployed, IAM-locked; DSAR live-verified 2026-08-10)

2.4 Cell 36 — intent-causal (VALIDATE/DECIDE)

Location: src/cells/cell36/causal_cell/ Key modules (3,043 lines total): - holdouts.py (300 lines) — holdout assignment, deterministic-identity-only gate - estimators.py (639 lines) — DR-Learner, CUPED estimators - refutation.py (333 lines) — arm-level permutation refutation battery - orchestrator.py (577 lines) — incrementality report generation - activation.py (112 lines) — uplift_export_cohort gated activation - storage.py (512 lines) — BigQuery storage, outcome ingestion, probabilistic identity exclusion

Causal capabilities: - ✅ Holdout assignment (deterministic hash, 10% permanent) - ✅ DR-Learner estimator - ✅ CUPED estimator - ✅ Refutation battery (arm-level permutation) - ✅ Probabilistic identity exclusion (3 enforcement points) - ✅ uplift_export_cohort → guardrails → DCP activation path - ❌ Reports held in-memory only (lost on restart; no BigQuery table) - ❌ No ghost-ad/bid-log integration - ❌ No caused-vs-anticipated credit classification yet

Lines: 4,705 (source ~3,060, tests ~1,640) | Tests: 91 functions | Dependencies: all exact pins

Status: PRODUCTION-READY (live-verified 2026-08-11)

2.5 Cell 37 — market-signal-ingest (SENSE / external intelligence)

Location: src/cells/cell37/ Lines: 1,309 (source only) | Tests: ZERO — no test directory exists | Dependencies: MIXED PINS (reproducibility concern per .claude/rules/03-dependency-reproducibility.md)

Cell 37 uses virtuoso_call indirectly via the /extract endpoint (delegates to virtuoso_models.extraction.market_extract.extract_market_signals). This is the only LII-adjacent cell with LLM calls.

Dependency concern: requirements.txt uses floor-only specs for most libraries (fastapi>=0.111, pydantic>=2.7, google-cloud-bigquery>=3.20, etc.) — only google-cloud-firestore, google-api-core, google-genai, and anthropic have upper bounds. This violates the dependency reproducibility rule (03-dependency-reproducibility.md, the Cell 03 GraphRAG cache incident).

2.7 Shared library — src/shared/mizoki_intent/

Key modules: - signal.py — signal validation, audio prohibition, consent evaluation, payload contract - taxonomy.py — taxonomy loader, deny-list enforcement, boot refusal - subject.py — GDPR subject access/erasure with count→delete→recount guarantee - client.py — IntentClient SDK, IntentScore.as_evidence()

2.8 SDK library — libs/mizoki_lii/ (4,089 lines)

Modules: client.py (890 lines), governance.py (439 lines), transport.py (405 lines), auth.py (167 lines), config.py (234 lines), errors.py (83 lines) Tests: 5 test modules (1,430 lines)

Provides LIIClient and LIIConfig for external consumers. Observe-only governance is enforced in the library.

2.9 MCP tool surface

mcp/mizoki-lii-mcp/ — 5 tools (observe-only):

Tool Shipped Boss equivalent Cell Endpoint
lii_score_intent intent_score_get 34 GET /v1/intent/score/{identity_id}
lii_list_inmarket intent_cohort_query 34 GET /v1/intent/cohort
lii_get_causal_credit intent_incrementality_report 36 GET /v1/intent/incrementality
lii_register_holdout intent_holdout_register 36 POST /v1/holdout/register
lii_subscribe intent_transitions_recent 34 GET /v1/intent/transitions

Boss Agent shipped tools (8 intent_* tools): intent_score_get, intent_cohort_query, intent_transitions_recent, intent_explain, intent_taxonomy_list, intent_consent_stats, intent_incrementality_report, intent_status. The 5 lii_* stdio tools are built but Boss registration is owner-gated (open-work item 19).


3. Causal / Lift-Engine Inventory

3.1 Cell 36 causal modules (3,043 lines)

Covered in §2.4 above.

3.2 Lift-engine service (services/lift-engine/) — 12,716 Python lines

Core modules:

Module Lines Capability
src/core/meta_learners.py 857 X-Learner, T-Learner, S-Learner meta-learning
src/core/causal_refutation.py 897 DoWhy refutation battery (placebo, random common cause, data subset)
src/core/causal_discovery.py 1,051 Structure learning (PC algorithm, GES, LiNGAM)
src/core/dynamic_uplift_rl.py 1,572 RL-based dynamic allocation (UCB, Thompson sampling)
src/core/interference.py 225 Network interference detection
src/core/event_study.py 104 Event-study designs
src/core/uplift_scoring.py 119 Uplift score computation
src/core/uplift_metrics.py 99 AUUC, Qini, cumulative gain
src/core/window_sizing.py 129 Experiment window sizing
src/core/allocation_sql.py 263 Budget allocation SQL generation
src/core/bayesian.py 305 Bayesian uplift estimation
src/core/recommender.py 357 Experiment-informed recommendations
src/core/lag_analysis.py 81 Lagged outcome analysis
src/evaluation/causal_benchmark.py 1,131 Causal benchmark suite
src/evaluation/counterfactual_evaluator.py 1,579 Counterfactual evaluation
src/api/causal_routes.py 449 Causal API routes
src/api/routes.py 308 General API routes
src/pbc/lift_measurement.py 196 Lift measurement framework

3.3 What the lift-engine CAN serve for the LII plan

LII Plan requirement Module Status
Caused-vs-anticipated credit classification services/measurement-rails/causal_credit.py (219 lines) BUILT. classify() splits treatment conversions: first N (baseline_rate × treatment_units) are "anticipated", remainder are "caused". v1.0.0 matched-baseline method. Deterministic identities only. Append-only bitemporal ledger. Consumed by Boss causal_credit_ledger_source.py → ROIEdge → budget reallocator.
Ghost-ad / holdout service services/measurement-rails/ghost_bid.py (326 lines) + Cell 36 holdouts.py (300 lines) + Cell 39 holdout.py (181 lines) BUILT. GhostBidExecutionService: shadow-bid/real-bid decision per unit. Holdout arm bids are WITHHELD and shadow-logged as counterfactuals. Pre-exposure gate. Arm provenance verification. All stores use SHA-256 deterministic bucketing, write-once.
Refutation framework Cell 36 refutation.py + lift-engine causal_refutation.py + Cell 26 DoWhy THREE parallel implementations. Cell 36 has arm-level permutation (no DoWhy dep, 333 lines); lift-engine has full DoWhy battery with sensitivity analysis (897 lines); Cell 26 has causalml+DoWhy (within 1,024-line Cell26.py). Need to decide which is canonical.
Causal discovery Lift-engine causal_discovery.py (1,051 lines) + Cell 23 (483 lines) Built in two places (PC, GES, LiNGAM, FCI, NOTEARS, hybrid ensemble). Not wired to intent data.
RL/bandit allocation dynamic_uplift_rl.py (1,572 lines) Built (Thompson Sampling, UCB, DynamicUpliftAgent, SequentialTreatmentOptimizer, CausalForestBandit, MarketingCampaignRL, OnlineCATEEstimator). Not wired to intent cohorts.
uplift_export_cohort Boss uplift_cohort_exporter.py (705 lines) + Cell 36 activation.py (112 lines) Built — the governed activation path. Boss exporter assembles top-K export with iROAS/iCPA reporting, guardrail validation, channel adapters (Google Ads, Meta, Email). Cell 36 builds the exact tool inputs.
Off-Policy Evaluation Lift-engine counterfactual_evaluator.py (1,579 lines) + services/replay-sim-ope/main.py (1,158 lines) Built in two places. Lift-engine: IPS, SNIPS, DR, DM. Replay-sim-ope: DR, WDR, MAGIC with bootstrap CIs and 4-check calibrated gate verdict (pass/soft_block/block) for Cloud Deploy pre-promote.

3.4 Other causal modules in the repo

Group Location Lines Purpose
Cells 26+27 src/cells/cell26/Cell26.py, cell27/Cell27.py ~2,858 X-Learner/DRLearner/SLearner (Cell 26); BayesianDecisionEngine, Thompson Sampling (Cell 27)
Cell 13 src/cells/cell13/causal_api.py 1,172 8 causal methods (PSM, IPW, doubly robust, DiD, IV, ATT)
Cell 23 src/cells/cell23/causal_discovery.py 483 PC, NOTEARS, FCI, hybrid ensemble
Cell 25 src/cells/cell25/causal_journey_integration.py 355 CATE reads from Cell 26 → KG edge weight updates
Cell 39 src/cells/cell39/outreach_cell/holdout.py, lift.py 413 CRE holdout/lift (ITT design, incremental lift CI)
measurement-rails services/measurement-rails/ ~1,094 Causal credit classification + ghost-bid (see above)
replay-sim-ope services/replay-sim-ope/main.py 1,158 OPE gate for Cloud Deploy pre-promote
relu-evaluation services/relu-evaluation-service/ ~1,433 UpliftScorer, ABAssigner, PolicyEvaluator
Boss agent miz-oki-adk-agents/boss/ causal modules 1,720 uplift_export_cohort (705), causal_credit_ledger (402), intent_api_integration (613)
Shared/misc Various (causal_edge_pacing, causal_journey_client, Cell 03 causal, etc.) ~2,595 Causal journey client, KG-based path finding, intervention simulation, budget reallocation

Total causal code in the platform: ~41,500 lines across all groups. The previously cited ~7,087 figure likely referred to lift-engine core only (src/core/ + src/evaluation/ = ~7,208 lines). The full inventory is nearly 6× larger. No causal module uses virtuoso_call — the entire causal codebase is pure statistical/ML (numpy, sklearn, scipy, econml, causalml, DoWhy).


4. Schema Gap Analysis

4.1 Canonical Event Envelope

Source: contracts/mizoki_contracts/envelope.py

The envelope payload field is Dict[str, Any] — an unconstrained dict. Schema enforcement happens at the cell level, not the envelope level.

4.2 IntentSignal payload contract (Cell 33)

Source: src/shared/mizoki_intent/signal.py

ALLOWED_PAYLOAD_KEYS = frozenset({
    "signal_type", "topic_id", "strength", "identity_kind", "consent",
})

This is a CLOSED contract. Any key not in this set is REJECTED (fail-closed). The behavioral micro-signals the LII plan requires are ALL missing:

Required field Status Notes
dwell_ms MISSING Not in payload contract; dwell_sec exists locally in Cell 32 only
scroll_velocity MISSING Not anywhere in canonical schemas
scroll_depth MISSING from envelope Cell 02 session_reconstructor.py tracks locally; Cell 02 BigQuery DDL has scroll_depth INT64
partial_watch_ratio MISSING Not anywhere
hover_duration MISSING Not anywhere
cart_abandon_velocity MISSING Not anywhere
search_refinement_count MISSING Not anywhere
cross_device_session_link MISSING Not anywhere
time_of_day_pattern MISSING from envelope time_of_day exists as a local feature in Cells 5/20/26
category_co_browse MISSING Not anywhere

Impact: Extending ALLOWED_PAYLOAD_KEYS requires a coordinated change across signal.py, BigQuery DDL (intent_signals table schema), the backtest harness, and the feature pipeline. This is the schema work the plan assumes exists.

4.3 JourneyEvent schema

Source: src/shared/virtuoso_models/schemas/journey-event.json

The context.ext field (additionalProperties: true) is the only extensibility point, but micro-signals are not defined there. The actor object carries user_id, email, phone_sha256, device_ifa, ip, ua — basic identity fields, no behavioral signals.

4.4 MarketSignal schema

Source: src/shared/virtuoso_models/schemas/market-signal.json

Market-level signals only (SERP position, trend index, sales rank, etc.). No person-level behavioral signals — by design (Cell 37 rejects person identifiers).


5. Graph Schema / Ontology Check

5.1 Cell 35 implemented graph model

Four node types, four edge types — frozen namespace (see §2.3 table).

Critical gaps for the LII plan's interest graph:

Required capability Exists? Notes
PRECEDES co-occurrence edges NO Described in OFFERING_MAP F.4 and blueprint §3 but not implemented
Household nodes/links NO Tests verify probabilistic household exclusion, but no household data model
Interest embeddings NO ADR-LII-001 explicitly confirms: "no vector index, vector store, or embedding of any kind"
Cross-device session links NO Not in graph model
Topic co-occurrence reasoning NO IntentTopic nodes exist but no inter-topic edges
Temporal sequence modeling NO No sequence-aware edges or ordering

5.2 Ontology registry

Source: ontology/registry.yaml v1.2.0

No interest, topic, co-occurrence, affinity, embedding, or vector node/edge types exist in the governed ontology. Cell 35's graph types (IntentIdentity, IntentTopic, etc.) are cell-local constructs with no ontology mapping.

Required OCPs for the LII plan: - Architecture-level: LatentIntent as a governed concept (requires owner approval) - Steward-level: InterestEmbedding, TopicCoOccurrence, HouseholdLink, BehavioralMicroSignal - Domain module: marketing/intent domain module (planned but not materialized)


6. API Contract Audit

6.1 Existing LII endpoints (Cell 34 — the canonical Intent API)

Method Path Purpose Status
GET /v1/intent/score/{identity_id} Single-identity intent score ✅ Live
POST /v1/intent/scores:batch Batch scoring (≤500/request) ✅ Live
GET /v1/intent/cohort In-market cohort query ✅ Live
GET /v1/intent/transitions Entered/left in-market feed ✅ Live
GET /v1/intent/explain/{identity_id}/{topic_id} Explanation paths via Cell 35 ✅ Live
GET /v1/intent/taxonomy Taxonomy registry ✅ Live
GET /health, /readyz, /metrics Observability ✅ Live
POST /a2a Agent-to-agent intent endpoint ✅ Live

6.2 GDPR/DSAR endpoints (all four cells)

Method Path Cell Purpose
GET /v1/intent/subject/{identity_id} 33, 34, 35, 36 Subject access (per-cell stores)
DELETE /v1/intent/subject/{identity_id} 33, 34, 35, 36 Subject erasure (per-cell stores)
POST /v1/intent/subject/{identity_id}:erase 34 only Cascade coordinator
GET /v1/intent/subject:readiness all four DSAR readiness probe

6.3 Cell 36 causal endpoints

Method Path Purpose
POST /v1/holdout/register Register holdout units
GET /v1/intent/incrementality Incrementality report
POST /v1/outcomes Outcome ingestion

6.4 Proposed endpoints from the LII plan — status

Proposed endpoint Exists? Notes
POST /lii/score NO (but GET /v1/intent/score/{identity_id} exists) Different path convention; Cell 34 already serves scoring
POST /causal/holdout Partial — POST /v1/holdout/register on Cell 36 Same capability, different path
GET /causal/credit/{conversion_id} NO No per-conversion credit lookup exists; Cell 36 serves aggregate incrementality reports only

6.5 Measurement-rails service (the causal credit backbone)

Service: measurement-rails (port 8045) | Auth: NO app-level auth — relies entirely on IAM posture | Status: NOT DEPLOYED by any CI workflow — operator RUNBOOK only | All flags shipped OFF

Method Path Purpose
POST /v1/rails/causal-credit:classify Per-conversion caused/anticipated classification against holdout baseline
POST /v1/rails/ghost-bid:execute Shadow-execute auctions under a registered holdout design
POST /v1/rails/designs Register measurement design (write-once, MDE required)
GET /v1/rails/designs List registered designs
POST /v1/rails/recompute:dry-run House re-attribution (in-memory)
POST /v1/rails/recompute House re-attribution (optionally persisted to BQ)
GET /v1/rails/drift/status Platform-vs-house drift streaks
POST /v1/rails/drift/record Feed one day per-source revenue
POST /v1/rails/writeback:recommend Recommend-mode value diffs (NEVER executes, autonomy L1)

Critical finding: measurement-rails is NOT in production/service-registry.yaml and has no CI deploy path. The "caused vs anticipated" credit classifier — a central requirement of the LII plan — lives here, not in the lift-engine.

6.6 Lift-engine service API surface

The lift-engine (services/lift-engine/) exposes routes via src/api/routes.py and src/api/causal_routes.py. These are a separate service from the LII cells, not yet wired to intent data.


7. Build Gates

Critical corrections required before any LII plan implementation can proceed. Numbered in severity order.

Gate 1 — Plan document must be sourced

Boris's anticipatory intent plan (LII_ANTICIPATORY_INTENT_PLAN_v1.md) is not in the repository. Without the canonical plan, build scope is undefined. Owner action: provide the plan document.

Gate 2 — Cell allocation requires Article VI owner directive

If the plan calls for new cells (40+), an Article VI directive is required (CONSTITUTION.md VI.1). The same-commit coherence obligation applies: CELL_REGISTRY, CONSTITUTION preamble, source_of_truth, byte-identical twins, check_conformance(), skillpack re-sync.

Owner decision needed: extend existing cells (33–36) vs. allocate new cells (40–42).

Gate 3 — IntentSignal payload contract must be extended

ALLOWED_PAYLOAD_KEYS is a 5-field frozen set. Adding behavioral micro-signals (dwell, scroll, watch ratio, etc.) requires a coordinated change across: - signal.py (add to ALLOWED_PAYLOAD_KEYS) - intent_bigquery_ddl.sql (add columns to intent_signals) - backtest/ harness (add to feature pipeline) - scoring_cell/scoring.py (extend compute_features() fv3) - Migration of existing data (new columns have no historical values)

This is the largest schema change in the LII plan and must be designed as an OCP before implementation.

Gate 4 — All LLM calls through virtuoso_call(Role.X)

Any new LLM-powered component (interest embedding generation, natural-language intent explanations, causal narrative generation) must route through virtuoso_call. No hardcoded model strings. The existing LII cells comply; new code must inherit this.

Gate 5 — Graph namespace frozen — extension requires explicit unlock

Cell 35's ALLOWED_LABELS and ALLOWED_EDGE_TYPES are frozen sets (graph_store.py:128-133). Adding PRECEDES, household edges, or interest embeddings requires: - Extending the frozen namespace (code change + tests) - OCP for the new node/edge types - Deciding whether the new types live in Cell 35 or a new cell

Gate 6 — No vector index/embedding store exists

ADR-LII-001 explicitly confirms zero vector/embedding infrastructure in Cells 33–36. The plan's two-tower embeddings and ANN retrieval (Cell 34 Phase B / Vertex AI Vector Search) are entirely unbuilt. If added, the DSAR erasure cascade MUST be extended to cover the new store — documented obligation in ADR-LII-001 §d.

Gate 7 — Label stream blocked (register item 6c)

The Phase A ranking gate is FAILED and OPEN. The model cannot be evaluated on real data because there are zero forward labels. Unblocking requires: - Shopify webhooks (orders/create) — the label stream - GA4 export link — session/page depth signals - First registered holdout (register item 23) — required for the third autonomy criterion (stable lift ≥ 2 purchase cycles)

No activation is legal without these. The LII plan's advanced capabilities compound this: more features require more signal, which requires more data sources.

Behavioral micro-signals are GDPR profiling (Article 22). The existing consent gate (evaluate_consent()) checks for intent_processing: True but the new signal types will carry richer personal data (browsing behavior, scroll patterns, dwell times). The consent basis and data-subject rights implications must be reviewed before any micro-signal is persisted.

Gate 9 — Identity resolution rules

The plan's household graph and cross-device linking introduce probabilistic identity resolution. Platform rule: probabilistic identities are excluded from causal math (3 enforcement points in Cell 36). The identity graph must: - Label every link deterministic or probabilistic with confidence - Enforce the deterministic-only gate on all causal paths - Never send person-level data to third-party enrichment (entity-spine scope)

Gate 10 — Dual refutation implementations

Cell 36 has arm-level permutation refutation (333 lines). The lift-engine has full DoWhy refutation (897 lines). The plan must decide which is canonical and whether to merge them. Running two parallel refutation batteries creates a consistency risk.

Gate 11 — Incrementality reports are ephemeral

Cell 36 holds reports in-memory only (bounded to MAX_REPORTS_KEPT = 100, lost on container restart). No BigQuery table exists for intent_incrementality_reports. The DDL defines no such table. Reports are not auditable after the fact. The plan's credit ledger requires durable storage.

Gate 12 — Observe-only is structural, not configurable

Nothing in the repository can promote LII out of observe-only automatically. The promotion criteria (Brier ≤ 0.20, AUC ≥ 0.72, stable lift ≥ 2 purchase cycles, human approval) are documented gates, not code thresholds. The plan must not assume any activation capability exists.


8. Existing Infrastructure — Surprises

Surprises (positive)

  1. Cell 36 is live-verified with full refutation. The causal-credit cell is the most production-ready component in the LII stack. The plan's incrementality requirements compose with it, not replace it.

  2. The lift-engine has 12,716 lines of causal code. Meta-learners (X/T/S), DoWhy refutation, causal discovery (PC/GES/LiNGAM), RL allocation (UCB/Thompson), Bayesian uplift, counterfactual evaluation — all built. The plan's causal requirements are largely a wiring exercise, not a build.

  3. The fv2 model crosses the Brier gate. The backtest report (2026-08-16) shows the fv2 logistic model achieves Brier 0.1845 mean (≤ 0.20 on 8/8 seeds). The fallback scorer does not. This is not yet the production path but the model quality work is ahead of what the plan likely assumes.

  4. DSAR is live-verified. The subject-rights cascade across all four cells has been end-to-end tested against real BigQuery, including the streaming-buffer 409 and clean 200 retry. Any new data store the plan adds must extend this cascade.

  5. The libs/mizoki_lii SDK exists (4,089 lines). External consumers have a governed client library. The plan's "democratization" goal already has a foundation.

Surprises (negative)

  1. PRECEDES edges are described in 3 documents but do not exist in code. The OFFERING_MAP, blueprint, and skillpack all describe co-occurrence reasoning via PRECEDES edges. Cell 35's frozen namespace does not include them and the code has never implemented them.

  2. Zero behavioral micro-signal fields in any canonical schema. The plan assumes rich behavioral data (dwell, scroll, hover, watch ratio) can flow through the platform. The payload contract is a 5-field frozen set. Even the JourneyEvent schema has no such fields.

  3. The ontology has no LII domain module. Cell 35's graph types are cell-local constructs with no governed ontology mapping. The plan's interest graph requires ontology work.

  4. Three refutation engines with no shared interface. Cell 36, the lift-engine, and Cell 26 each have their own refutation implementation. This must be rationalized.

  5. Cell 37 has zero tests. The only LII-adjacent cell with no test directory. Given its virtuoso_call usage (the /extract endpoint), this is a coverage gap.

  6. Cell 37 dependency pins are floor-only. Most requirements use >= without upper bounds, violating the reproducibility rule derived from the Cell 03 GraphRAG cache incident (03-dependency-reproducibility.md).

  7. measurement-rails is not in the service registry and has no CI deploy path. The "caused vs anticipated" credit classifier — central to the LII plan's causal credit requirement — lives in this unregistered service.

  8. The boss_agent_skillpack.md quick-reference still references the never-built POST api.mizoki.ai/api/v1/causal/holdout at line 169 — contradicting the documented correction at line 35 of the same file (corrected 2026-08-11).

  9. Service registry says cells 33–36 are "designed" but they are deployed and serving (Cell 36 live-verified). The registry is stale.


9. Owner Decisions Needed

# Decision Impact Precedent
D1 Source the plan document (LII_ANTICIPATORY_INTENT_PLAN_v1.md) Blocks all build work —
D2 Extend existing cells (33–36) vs. allocate new cells (40–42) Article VI directive if new cells CRE ruling: extend inline (ADR-LII-001 (a))
D3 Behavioral micro-signal schema design (OCP) Gate 3 — largest schema change Closed payload contract in Cell 33
D4 Canonical refutation engine (Cell 36 vs. lift-engine) Gate 10 —
D5 Graph namespace extension plan (which types, which cell) Gate 5 Frozen namespace in Cell 35
D6 Incrementality report durability (BigQuery table) Gate 11 Cell 36 in-memory only
D7 Data source prioritization (Shopify → GA4 → Klaviyo → TTD) Gate 7 — label stream Register item 6c
D8 Consent basis for behavioral profiling signals Gate 8 — GDPR Art 22 Existing intent_processing consent

10. Phase Summary

Dimension Finding
Plan document NOT FOUND — must be sourced
Existing LII cells (33–37) All four core cells deployed and IAM-locked (~43,749 lines, 1,023 tests). Cell 36 live-verified. Cell 37 partial (zero tests).
Causal code ~41,500 lines across 20+ modules (lift-engine 12,716 + Cell 36 4,705 + measurement-rails 1,094 + Cells 26/27 2,858 + scattered). Rich but largely unwired to intent.
Next free cells 40 onward (38–39 = CRE)
Behavioral micro-signals Zero fields in any canonical schema — the largest gap
Interest graph / co-occurrence Not implemented. PRECEDES described but absent.
Vector/embedding store Does not exist. Confirmed by ADR-LII-001.
Label stream Blocked — zero forward labels (register item 6c)
Consent model Exists but needs extension for micro-signal profiling
Boss MCP tools 8 shipped intent_* tools; 5 lii_* tools built, owner-gated
Observe-only Structural — no code path to promote
DSAR Live-verified — any new store must extend cascade

Report generated 2026-08-18 as Phase 0 of the LII Anticipatory Intent Integration. No code was written. Next phase requires the plan document and owner decisions D1–D8.

← All docsView source on GitHub →