Intent API Platform — Phase A Gate Report

Date: 2026-08-01 · Status: SHADOW (unchanged) · claim_label: built, pre-benchmark Corpus: docs/INTENT_CORPUS_CARD_MYCOCOONS.md · Plan of record: docs/INTENT_API_PLATFORM_BLUEPRINT.md §5/§9 Named-human sign-off: ______ (required for any Phase A exit decision)

Every figure below is a verified result measured on live infrastructure unless labeled otherwise.

Phase A exit gates — 5 of 7 closed, 1 failed, 1 pending

Gate State Evidence
Leakage regression test green ✅ CI + structural (available_to_model_at <= label_cutoff in every feature read; offline twin tested)
Consent fail-closed negatives ✅ CI + live (absent consent → denied; EU non-explicit → denied; denial rows carry {region, reason, source_system} only)
Deny-list / audio negatives ✅ CI + live (audio_ambient → rejected; political_donors → rejected; sensitive read → 422)
Zero consumers (Shadow posture) ✅ boss has no INTENT_* env; UI has no INTENT_API_BACKEND_URL
Calibration ECE recorded on time-split eval ✅ recorded ECE 0.0006, Brier 0.0006 (fv2, calibrated, deduplicated eval — most-recent-20%-of-cutoffs slice, 104,782 rows / 61 positives). Caveat stated: at a ~0.06% eval base rate these are near-trivially small — calibration is honest but weakly informative
AUUC vs naive-recency baseline ❌ FAILED Model cumulative-gain area 0.5703 vs naive-recency 0.7492 (must beat baseline; clean deduplicated eval). Three attempts, protocol held constant — see below
Cost actuals vs §7 targets ⏳ Needs a billing window; services have run ~1 day

The ranking-gate result (the honest headline)

On the only real corpus available, a boosted-tree intent model does not beat a naive recency sort at predicting 30-day repeat purchase:

Attempt Config Eval AUC Model gain area Recency gain area
1 fv1 (5 features), unweighted 0.5749 0.5789 † 0.7414 †
2 fv1, auto_class_weights 0.5749 0.5454 † 0.7329 †
3 fv2 (9 RFM features), weighted 0.5749 0.5771 † 0.7436 †
3 (clean) fv2, weighted, deduplicated eval 0.5749 0.5703 0.7492

† Attempts 1–3 as first computed were contaminated by the calibration range-join row multiplication (defect 7 below) — eval row counts were inflated ~3×. The clean deduplicated re-measurement (104,782 rows, 61 positives) confirms the same conclusion with the same protocol.

Protocol: identical corpus (383,582 examples, 655 positives — 0.17%), identical time split (most recent 20% of weekly cutoffs 2022-01 → 2025-09), identical baseline (rank by days_since_last_signal ascending). Model verified to carry all 9 fv2 features (ML.FEATURE_INFO); training early-stopped at 6 iterations.

Interpretation (assessment, not a measurement): the corpus is order-anchored — the only behavioral signals are prior purchases and their acquisition context. In that information regime, "bought recently" is close to the whole signal, and a tree's piecewise-constant scores tie large groups that a continuous recency sort orders finely. This is a data limitation, not a platform defect: with real pre-purchase clickstream (absent from the project — no GA4 raw export exists) the model would have signals recency cannot see.

Not done, deliberately: no further metric-chasing iterations, no synthetic positives, no baseline weakening, no cherry-picked horizon. The gate stays open until either (a) richer real signals flow (live traffic into Cell 33), or (b) the named human amends the gate for this corpus class.

What is now running (all live-verified)

Defects found and fixed during this rollout (all committed)

  1. INFORMATION_SCHEMA.MODELS does not exist in BigQuery — aborted the procedure install; replaced with an exception-probed ML.TRAINING_INFO.
  2. Cell 33's DATA-001 first-ingest stamping silently emptied every historical cutoff — resolved with the documented retrospective availability reconstruction (corpus card).
  3. 20GB maximum_bytes_billed cap aborted BQML training — train/calibrate now run uncapped.
  4. Bulk-patcher substring splice duplicated a CAST block → BigQuery duplicate column; deduplicated (the Edit replace_all hazard's SQL cousin).
  5. Cell 34 cloudbuild --set-env-vars (replace semantics) wiped the operator-set ALLOWED_CALLER_SA → fail-closed 503s; switched to --update-env-vars and restored the allowlist.
  6. run_live_scoring N+1: per-identity prev-state queries hung 5k identities past the request deadline; prev-state now joins inside the prediction query and scores batch-insert.
  7. Calibration range-join row multiplication: tied raw scores give NTILE bins identical boundaries; the range join matched several bins per prediction — 2,694 pairs scored as 8,694, and eval rows inflated ~3×. Fixed with a deterministic MAX over matched bins in all four consumers; live-verified scored == pairs == 2,694 after the fix.

Operator side-findings (not intent work; surfaced by the stream survey)

Live extenders (added 2026-08-01, evening — operator-directed)

Both recommended streams are now BUILT AND DEPLOYED, fail-closed until two operator steps (docs/INTENT_EXTENDERS_RUNBOOK.md):

Service Posture Activates when
intent-shopify-extender intended-public (HMAC is the auth — Shopify cannot mint OIDC; secret currently random-rotated placeholder → every webhook 401s) operator creates the store webhooks + pastes the real signing secret
intent-ga4-extender IAM-locked; hourly scheduler intent-ga4-projection live; /run answers export_not_configured no-op operator links GA4 → BigQuery export + sets GA4_EXPORT_DATASET (the ekis-ga4 secret was measured to be a placeholder template — no real property is wired anywhere)

Supporting changes, all live-verified: Cell 33 gained POST /v1/outcomes (the live ground-truth door — consent fail-closed + counted, deny-list 422, bitemporal floor with provenance.backfill, idempotent; accept + denial both proven on rev intent-signal-ingest-00005-ss2); intent-extender-sa created (least-privilege: run.invoker on Cell 33 + bigquery.jobUser) and added to Cell 33's allowlist. Fleet governance note: the intended-public register grows 11 → 12 (intent-shopify-extender).

Recommendation to the named human

Keep Phase A open in Shadow (costs are minimal; hourly scoring exercises the whole loop). Complete the two runbook steps to start the live corpus: (1) Shopify webhooks + real signing secret, (2) GA4 BigQuery export link + GA4_EXPORT_DATASET. Set GA4 user_id to the Shopify customer id on the storefront so browse signals join purchase outcomes deterministically. Once real signals + outcomes accumulate, re-run the gate eval — the ranking gate gets its honest retry with signals recency cannot see.

← All docsView source on GitHub →