ORACLE Pre-Conversion Intent v1.0 — STEP 0 Verify-Before-Build Report
Date: 2026-08-18
Session: oracle-preconv (cloud; master prompt v1.2 — confirmed running v1.2 after the owner supplied it; v1.1 executed first, v1.2 patches folded in)
Plan of record: docs/ORACLE_PRECONVERSION_INTENT_v1.0.md — sha256 bd8096f786d0e31d356f82ab30d0d26266d93c71dbf093cb4f296ff31dd3fa66 (20,774 bytes, byte-identical to the Drive prompt-board copy 1OM0ZhFuEozRTgkzXIE_EHLzjKAFZFQe5)
Authority order applied: WIRING.md > OFFERING_MAP v2.2 > .claude/rules/03-canonical-architecture.md > plan of record > master prompt.
Rule applied throughout: where the registry/tree disagrees with the plan, the registry wins — the delta is recorded here and bindings adjusted; the plan is not forced.
claim_label for everything in this lane: built, pre-benchmark. Nothing below is a performance claim.
0. Coordination preconditions
| Check | Result |
|---|---|
| Ledger overlap | None. Session board + CLAUDE.md inbox + .claude/memory/inbox/2026-08.md read before build; no live claim touches cells 33–35 / creative vectors / bridges. The open cell37-external-kg-v1.2-build claim's src/shared/virtuoso_models overlap is landed history (PR #716 merged, d22bb99). |
| Shared memory | shared-memory/memory.py recall ×2 (oracle preconversion / cell33 behavioral schema) = 0 matches — no prior or duplicate build. |
| Claim recorded | claim: oracle-preconversion-intent-v1.0 recorded 2026-08-18T20:10:54Z, landed on origin/main (memory-only commit 5bbf95e, ancestry-verified; survived the e929cda theirs-merge intact — re-grepped on origin/main). Session-board row added. |
| Branch discipline | Build branch claude/oracle-preconversion-intent-v1.0 — local only, no push until in-session APPROVED: MERGE (GATE 1). The only pushes so far are memory/coordination commits on claude/new-session-b63rxg (claim visibility per rule 02; precedent: the cell37 lane's GATE-1 claim landed the same way). |
0b. v1.2 patch — standing-gate script verification (FALSE-GREEN FOUND)
The v1.2 patch mandates verifying each standing gate's actual filename and invocation, because "a gate that errors on a missing script is a FAILED gate, never a passed one." Verifying surfaced a false green in v1.1's gate list:
python3 scripts/mizoki_canon.py --checkis a NO-OP that always exits 0.scripts/mizoki_canon.pyis a library — noargparse, no__main__, no--checkhandling (grep-verified; its own docstring names the real consumers). Running it with--checkimports the module and exits 0 without checking anything. This matches the y3m3cq pre-flight review's finding F1 (docs/reports/ORACLE_PRECONV_PROMPT_REVIEW_2026-08-18.md, HIGH). Any prior "canon clean via mizoki_canon --check" claim is void.- The verified standing-gate commands, used at every gate hereafter:
1. Full pytest — per cell, fresh venv.
2. Canon lint:
python3 scripts/skill_sync.py --audit(exit 1 on HIGH; the real canon gate — replaces the no-op). 3. Skill parity:python3 scripts/skill_sync.py --checkandpython3 scripts/skills_sync.py --check(BOTH exist — the first checks 14 skills + derived copies + registry; the second checks skill-copy byte-identity/version parity) andpython3 scripts/ontology_skills_sync.py --check. 4. Hygiene greps (rule 03 V1 SRDAL / V2 retired-model / V3 PA-API). 5.python3 scripts/claude_memory.py check --strict. - Measured results on the rebased branch:
skill_sync.py --auditexit 0 (no HIGH; only pre-existing LOW findings in skill files this build never touched),skill_sync.py --check/skills_sync.py --check/ontology_skills_sync.py --checkall exit 0, hygiene greps clean on changed source, memory strict valid.
0c. v1.2 patch — Cell 37 canon delta (recorded)
The v1.2 patch requires recording Cell 37's actual registry identity vs OFFERING_MAP v2.2 and queueing an OFFERING_MAP correction if it drifts. Measured:
- OFFERING_MAP v2.2 names Cell 37 "Data Injector". The deployed Cloud Run service is
market-signal-ingest(src/cells/cell37/).docs/architecture/CELL_REGISTRY.mdrow 37 (the number authority) already reconciles both: namemarket-signal-ingest, canonical role "Data Injector & External Intelligence Gateway" (owner ruling 2026-08-11), external-intelligence scope largely[IN BUILD], with an explicit "stated role vs. built reality" note. - Assessment: this is already-recorded canon, not new unrecorded drift. CELL_REGISTRY.md (which outranks OFFERING_MAP on cell numbers/identity) already carries the reconciliation with its evidence. This build uses Cell 37 only as external-evidence corroboration for latent bridges (Workstream D), which matches the canonical role exactly.
- OFFERING_MAP correction queued (end-of-run report): OFFERING_MAP v2.2's bare "Data Injector" label for Cell 37 could be tightened to match CELL_REGISTRY.md's fuller "Data Injector & External Intelligence Gateway (
market-signal-ingest; external-intel scope [IN BUILD])". This is a docs-only tightening for the canon-owner; this build does not touch OFFERING_MAP. Recorded so the delta does not survive this run unrecorded.
1. Live Cell Registry — ground truth
gcloud is not installed in this cloud sandbox (which gcloud = exit 1; rule 02 records this as a known environment fact). The Cloud-Run half of the registry check is therefore not runnable from this session; the registry files + live-state records are the ground truth used, and every claim below about serving state is a repo-recorded state, not a fresh probe:
docs/architecture/CELL_REGISTRY.md(number authority, 39 registered cells): cells 33–36 =intent-signal-ingest/intent-scoring-api/intent-graph/intent-causal— all deployed, IAM-locked, dispatch-only (deploy-intent-platform.yml); cell 36 live-verified 2026-08-11 (run 31494599562). Cell 34 rev00018-c7bserving per lii-deploy-smoke row (2026-08-12).- Cell 37
market-signal-ingest(Data Injector & External Intelligence Gateway): deployed, partially operational; latest revision00011-tfmper register item 25; extract leg not green (operator-held). Its MarketSignal layer (market_signal.py,market_ingest_gate) is the external-evidence path Workstream D treats as corroboration only. - Plan deltas (registry wins): the plan's header says "37-cell platform" (registry: 39); its C1 table calls Cell 35 "Neo4j intent graph" (registry: Firestore-journal + in-memory serving; Neo4j retired 2026-08-09,
neo4j-uriNXDOMAIN by choice) and Cell 37 "data-injector" (registry name:market-signal-ingest, role per owner ruling 2026-08-11).
2. JourneyEvent schema + response_schema_hash baseline
- Schema:
src/shared/virtuoso_models/schemas/journey-event.json(byte-identical vendored twin underservices/virtuoso-models-service/virtuoso_models/). Nobehavioral_signalblock exists today — plan delta recorded; Workstream A adds it additively. response_schema_hash=sha256(canonical(journey-event.json)), computed live injourney_event.schema_hash()and auto-stamped into provenance ("never hand-write" — package CLAUDE.md). Baseline hash at Step 0, measured in-session (python3 -c "…schema_hash()"on this tree):686738a05b369f325ca9e2924a622a44db6575ebb66aa6f5b2122860a2cd9161(corroborated: matches the686738a05b36schema pin recorded on the session board by the cre-outreach-p3 lane's mapper union merge)- The bump mechanism is therefore: change the schema file additively → the stamped hash moves everywhere automatically. No test pins a literal hash (both suites assert
stamped == computed), so the deliberate bump needs no pin edits; the build adds an explicit old-events-still-validate test (rollback level 3). - Two envelopes, both governed, neither bypassed: cell33's IntentSignal is a constrained projection of
mizoki_contracts.envelope.CanonicalEventEnvelope(the R2 "no parallel schema" ruling) — that is the consent-gated intent door. The virtuoso_models JourneyEvent is the Sense-phase journey layer (MAPPERS →ingest_gate, MERGE onevent_id, CAS onsource_payload_hash). The plan's A.3 language maps to the JourneyEvent side; the behavioral extension lands on both, additively, with the same rejection gates.
3. Cell 34 model lane
| Plan says | Tree says (wins) |
|---|---|
| "matrix factorization + boosted tree, nightly batch" | BQML BOOSTED_TREE_CLASSIFIER (model/bqml.sql:105, intent_bqml_v1), hourly single-writer batch (workflows/intent_scoring_schedule.yaml, cron 7 * * * *), plus an offline fv2 logistic surrogate (backtest/model.py, fv2-logit-1) and a memory-mode heuristic scorer. No realtime inference path exists. |
"existing LII_REALTIME feature flag" |
Does not exist (repo grep: doc/skill prose only). Per Step 0 instruction it is created in this build: per-cell env flag, default OFF, source-literal asserted by test. No central flag registry is used by cells 33–36 (family convention = per-cell Settings + env; src/shared/flags.py is legacy-cell-only). |
| "circuit breaker unchanged: on outage, SDK serves cached stale=true" | No circuit breaker exists in cell34 (grep = 0 hits). The analogous machinery is readiness verdicts (ready|degraded|blocked) and Cell 33's freshness verdict gate. Honored by architecture instead: the transformer path is additive beside the batch path; any transformer-path failure falls back to the existing serving behavior. Nothing existing is modified. |
"LEARN loop writes realized/realized_at to unified.intent_predictions" |
Neither the table nor the columns exist in code (repo-wide grep = skill/doc prose only). The real label loop: mizoki_intent.intent_outcomes joined at horizon → sp_build_training_examples → mizoki_intent.intent_training_examples (leakage gate available_to_model_at <= label_cutoff, scoring.py:118-134). The transformer's label stream binds to this loop. |
"128-dim matching unified.intent_vectors" |
unified.intent_vectors does not exist (docs only). 128 stands as the design dimension for the new surfaces. Spec-level unified.* names resolve to dataset mizoki_unified_data (ADR-NY-001 §4, recorded in bigquery/schemas/measurement_rails_ddl.sql header). |
shadow table "unified.intent_predictions_shadow" |
Bound to the real scores table's shape instead: mizoki_intent.intent_scores_shadow (shadow twin of intent_scores, additive DDL, operator-applied). |
fv2 lockstep constraint honored: the five fv2 homes (scoring.py offline twin, bqml.sql, bqml_procedures.sql, ScoreStore.bqml_predictions, pipeline FEATURE_COLS) are not modified — the transformer is a separate, flag-off lane. Dependency note: cell34's image has no numpy/sklearn/torch; the transformer adds numpy with an exact pin (pure-CPU inference; latency benchmark test included).
4. Shopify Web Pixel (P1) — capture surface
Measured state (this repo):
- The pixel JS artifact is not in this repository — it ships with the Shopify app project (docs/reports/SIGNAL_SHOPIFY_LANE_STATUS_2026-08-12_R3.md:68-71, docs/roadmap/P1_BUILD_PLAN.md:92-95).
- The server side that exists here: webPixelCreate activation (services/service-marketing-connectors/shopify_install.py:651-737, settings allowlist {ingest_url, shop_domain}) pointing at {SHOPIFY_PIXEL_EXTENDER_URL}/pixel/events (main.py:838), deferring loudly while unset.
- /pixel/events does not exist yet on intent-shopify-extender (routes today: /v1/downstream/shopify, /health, /readyz). The P1 record says the receiving endpoint "is designed together with that artifact".
Binding for Workstream A (no second collector): this build ships (1) the extender's /pixel/events receiving endpoint — flag-off, consent-gated, forwarding into the single cell33 door — and (2) the shared batching/VDI capture core as a versioned JS artifact + settings contract in this repo, which the Shopify-app pixel wraps (that wrap is [IN BUILD] in the app project, out of this repo's reach). No parallel ingest path is created; the endpoint fans into POST /v1/signals exactly as the webhook extender does.
5. Consent gate + GDPR erasure cascade suites (the matrices the new fields join)
- Consent:
evaluate_consent(src/shared/mizoki_intent/signal.py:103) — fail-closed, affirmative purpose-specificintent_processing: true, EU/EEA/UK + unknown regions require explicit basis. Gate order in cell33 (orchestrator.py:230-246): audio → consent (denied = dropped+counted, never persisted) → envelope → look-ahead floor → closed-contract validation → idempotent persist. Scope form ("analytics-only does not qualify") lives inlibs/mizoki_lii/governance.py(CONSENT_SCOPE="intent_modeling") with both-direction tests. - Consent suites:
src/cells/cell33/tests/test_ingest_api.py(ConsentGateTests/AbsentConsentGateTests + region sweep),test_outcomes_api.py,libs/mizoki_lii/tests/test_governance.py, extendertest_extender.py::test_no_consent_never_forwarded. - Erasure: registry
IDENTITY_LINKED_TABLES(cell33storage.py:493— 6 tables) minusERASURE_EXCLUDED_TABLESmust equal the union of every cell'sSUBJECT_STORES— enforced by the divergence guardsrc/cells/cell34/tests/test_erasure_coverage.py. Cascade coordinator: cell34POST /v1/intent/subject/{id}:erase. Graph leg: cell35DELETE /v1/identity/{id}(memory sweep + Firestore journalsubject_keystamps;SUBJECT_OWNED_LABELS, count→erase→recount receipts). - Erasure suites the new surfaces join:
cell33/tests/test_subject_rights.py,cell34/tests/test_subject_rights.py+test_erasure_coverage.py,cell35/tests/test_graph_erasure.py,test_erasure_api.py,test_erasure_cascade.py,test_subject_rights.py,cell36/tests/test_subject_rights.py, plus the fourtest_dsar_readiness.py. - Deny-list: dual enforcement (write-time 422 in
validate_signal_payload; boot refusal inload_taxonomy; read-path scrubscrub_deny_listed).BASE_DENY_TERMSin code can only be extended by YAML, never shrunk.
6. Plan-of-record artifact
docs/ORACLE_PRECONVERSION_INTENT_v1.0.md committed in this branch, byte-identical to the Drive prompt-board original (size match 20,774 B; base64-faithful transfer; sha256 above recorded in the coordination claim as well).
7. Cell 35 graph namespace — Workstreams C/D binding
Tree facts: namespace is frozen allowlists (graph_store.py:128-157) — labels {IntentIdentity, IntentTopic, SignalAggregate, IntentExplanation}, edges {HAS_INTENT, HAS_EXPLANATION, CITES, EXHIBITED}, composite tenant-first identity keys, mandatory edge metadata {confidence, source, timestamp, tenant_id, provenance}, durable-journal subject stamps (SUBJECT_OWNED_LABELS, edge_envelope endpoint rule). There are no Topic nodes and no PRECEDES edges (plan delta; the shipped co-occurrence substrate is HAS_INTENT edges + intent_transitions).
Binding: the plan's (:Topic) → IntentTopic; (:Customer) → IntentIdentity (composite key). New labels Creative (key (tenant_id, creative_id)) and LatentBridge (key bridge_id = tenant-embedded hash, brg_… per the agg_/exp_ convention); new edges RESONATED_WITH (IntentIdentity→Creative), IN_STATE (IntentIdentity→LatentBridge), EVIDENCES (IntentTopic→LatentBridge). All three new-edge classes are subject-owned where they start at IntentIdentity and join the erasure sweep + journal stamps + subject snapshot; Creative/LatentBridge nodes are tenant-scoped, identity-free (aggregate class — deliberately out of erasure's reach, same rule as SignalAggregate, published to auditors via the retained-nodes list).
8. Model roles
Role.CREATIVE_MM (→ gpt-5.6-sol) and Role.DATA_CAUSAL (→ gemini-3.6-flash, GA flag on) both exist in model_registry.py — no registry changes needed; all LLM access in this build goes through virtuoso_call(Role.…); zero hardcoded model strings.
9. Out of scope (restated, binding)
No activation of any kind (hesitation retargeting, aesthetic serving, bridge targeting), no holdout registration, no dashboard surfacing (O-5 open), no change to causal-credit math, no deploys, no DDL application (operator-applied per house convention), no site-visible files. All new capability flags default OFF with source-literal tests.