Intent API Platform — Phase A Gate Report
Date: 2026-08-01 · Status: SHADOW (unchanged) · claim_label: built, pre-benchmark
Corpus: docs/INTENT_CORPUS_CARD_MYCOCOONS.md · Plan of record: docs/INTENT_API_PLATFORM_BLUEPRINT.md §5/§9
Named-human sign-off: ______ (required for any Phase A exit decision)
Every figure below is a verified result measured on live infrastructure unless labeled otherwise.
Phase A exit gates — 5 of 7 closed, 1 failed, 1 pending
| Gate | State | Evidence |
|---|---|---|
| Leakage regression test green | ✅ | CI + structural (available_to_model_at <= label_cutoff in every feature read; offline twin tested) |
| Consent fail-closed negatives | ✅ | CI + live (absent consent → denied; EU non-explicit → denied; denial rows carry {region, reason, source_system} only) |
| Deny-list / audio negatives | ✅ | CI + live (audio_ambient → rejected; political_donors → rejected; sensitive read → 422) |
| Zero consumers (Shadow posture) | ✅ | boss has no INTENT_* env; UI has no INTENT_API_BACKEND_URL |
| Calibration ECE recorded on time-split eval | ✅ recorded | ECE 0.0006, Brier 0.0006 (fv2, calibrated, deduplicated eval — most-recent-20%-of-cutoffs slice, 104,782 rows / 61 positives). Caveat stated: at a ~0.06% eval base rate these are near-trivially small — calibration is honest but weakly informative |
| AUUC vs naive-recency baseline | ❌ FAILED | Model cumulative-gain area 0.5703 vs naive-recency 0.7492 (must beat baseline; clean deduplicated eval). Three attempts, protocol held constant — see below |
| Cost actuals vs §7 targets | ⏳ | Needs a billing window; services have run ~1 day |
The ranking-gate result (the honest headline)
On the only real corpus available, a boosted-tree intent model does not beat a naive recency sort at predicting 30-day repeat purchase:
| Attempt | Config | Eval AUC | Model gain area | Recency gain area |
|---|---|---|---|---|
| 1 | fv1 (5 features), unweighted | 0.5749 | 0.5789 † | 0.7414 † |
| 2 | fv1, auto_class_weights |
0.5749 | 0.5454 † | 0.7329 † |
| 3 | fv2 (9 RFM features), weighted | 0.5749 | 0.5771 † | 0.7436 † |
| 3 (clean) | fv2, weighted, deduplicated eval | 0.5749 | 0.5703 | 0.7492 |
† Attempts 1–3 as first computed were contaminated by the calibration range-join row multiplication (defect 7 below) — eval row counts were inflated ~3×. The clean deduplicated re-measurement (104,782 rows, 61 positives) confirms the same conclusion with the same protocol.
Protocol: identical corpus (383,582 examples, 655 positives — 0.17%), identical
time split (most recent 20% of weekly cutoffs 2022-01 → 2025-09), identical
baseline (rank by days_since_last_signal ascending). Model verified to carry
all 9 fv2 features (ML.FEATURE_INFO); training early-stopped at 6 iterations.
Interpretation (assessment, not a measurement): the corpus is order-anchored — the only behavioral signals are prior purchases and their acquisition context. In that information regime, "bought recently" is close to the whole signal, and a tree's piecewise-constant scores tie large groups that a continuous recency sort orders finely. This is a data limitation, not a platform defect: with real pre-purchase clickstream (absent from the project — no GA4 raw export exists) the model would have signals recency cannot see.
Not done, deliberately: no further metric-chasing iterations, no synthetic positives, no baseline weakening, no cherry-picked horizon. The gate stays open until either (a) richer real signals flow (live traffic into Cell 33), or (b) the named human amends the gate for this corpus class.
What is now running (all live-verified)
- Corpus: 5,748 consent-gated signals + 2,877 outcomes from the My Cocoons Shopify store (5,062 gate-passing identities; 2,552 EU/UK/unknown excluded fail-closed; non-opted-in identities never left BigQuery).
- BQML pipeline as reviewed stored procedures (
sp_build_training_examples/sp_train_model/sp_build_calibration/sp_score_hourly/sp_emit_transitions) — degenerate-corpus train refusal and no-model score refusal both live-proven. - Model:
intent_bqml_v1(fv2, class-weighted) + isotonic calibration table — serving scores labeledbuilt, pre-benchmarkin Shadow. - Hourly automation (A.5): Cloud Scheduler
intent-hourly-scoring(cron7 * * * *UTC) → OIDC (runner SA, origin audience) → Cell 34POST /v1/internal/score-run-live→ BQML predictions → hysteresis → transitions → outbox → Pub/Sub → Cell 35 graph. - Scoring loop remains SHADOW: scores and transitions flow; no consumer reads them; every payload carries the claim label.
Defects found and fixed during this rollout (all committed)
INFORMATION_SCHEMA.MODELSdoes not exist in BigQuery — aborted the procedure install; replaced with an exception-probedML.TRAINING_INFO.- Cell 33's DATA-001 first-ingest stamping silently emptied every historical cutoff — resolved with the documented retrospective availability reconstruction (corpus card).
- 20GB
maximum_bytes_billedcap aborted BQML training — train/calibrate now run uncapped. - Bulk-patcher substring splice duplicated a CAST block → BigQuery duplicate
column; deduplicated (the
Edit replace_allhazard's SQL cousin). - Cell 34 cloudbuild
--set-env-vars(replace semantics) wiped the operator-setALLOWED_CALLER_SA→ fail-closed 503s; switched to--update-env-varsand restored the allowlist. run_live_scoringN+1: per-identity prev-state queries hung 5k identities past the request deadline; prev-state now joins inside the prediction query and scores batch-insert.- Calibration range-join row multiplication: tied raw scores give NTILE bins identical boundaries; the range join matched several bins per prediction — 2,694 pairs scored as 8,694, and eval rows inflated ~3×. Fixed with a deterministic MAX over matched bins in all four consumers; live-verified scored == pairs == 2,694 after the fix.
Operator side-findings (not intent work; surfaced by the stream survey)
cell3-profile-ingestionscheduler 422s every 15 minutes.gemini-meta-workerfails daily (known missing_META_AD_ACCOUNT_ID).- Schedulers
nightly-attribution-recomputeandcausal-edge-pacing-dailytarget nonexistent services. - Pub/Sub topic
mizoki-eventshas zero subscribers — anything published there is dropped.
Live extenders (added 2026-08-01, evening — operator-directed)
Both recommended streams are now BUILT AND DEPLOYED, fail-closed until two
operator steps (docs/INTENT_EXTENDERS_RUNBOOK.md):
| Service | Posture | Activates when |
|---|---|---|
intent-shopify-extender |
intended-public (HMAC is the auth — Shopify cannot mint OIDC; secret currently random-rotated placeholder → every webhook 401s) | operator creates the store webhooks + pastes the real signing secret |
intent-ga4-extender |
IAM-locked; hourly scheduler intent-ga4-projection live; /run answers export_not_configured no-op |
operator links GA4 → BigQuery export + sets GA4_EXPORT_DATASET (the ekis-ga4 secret was measured to be a placeholder template — no real property is wired anywhere) |
Supporting changes, all live-verified: Cell 33 gained POST /v1/outcomes
(the live ground-truth door — consent fail-closed + counted, deny-list 422,
bitemporal floor with provenance.backfill, idempotent; accept + denial both
proven on rev intent-signal-ingest-00005-ss2); intent-extender-sa created
(least-privilege: run.invoker on Cell 33 + bigquery.jobUser) and added to
Cell 33's allowlist. Fleet governance note: the intended-public register
grows 11 → 12 (intent-shopify-extender).
Recommendation to the named human
Keep Phase A open in Shadow (costs are minimal; hourly scoring exercises the
whole loop). Complete the two runbook steps to start the live corpus:
(1) Shopify webhooks + real signing secret, (2) GA4 BigQuery export link +
GA4_EXPORT_DATASET. Set GA4 user_id to the Shopify customer id on the
storefront so browse signals join purchase outcomes deterministically. Once
real signals + outcomes accumulate, re-run the gate eval — the ranking gate
gets its honest retry with signals recency cannot see.