LII backtest model quality — why 0.6884 was the wrong number twice

Date: 2026-08-16 Cell: 34 (intent-scoring-api) Register item: 17 — "LII offline backtest NO-GO" Claim label: built, pre-benchmark (unchanged — a synthetic corpus upgrades nothing) Reproduce:

python3 -m src.cells.cell34.backtest --panel

Summary

Register item 17 recorded the LII offline backtest at AUC 0.6884 against a 0.70 line, and against the 0.72 autonomy gate, and three further documents repeated it. The number is arithmetically correct and was honestly produced. It is nonetheless the wrong number to have been reading, for two independent reasons — and the second one is the important one.

  1. It measured the fallback, not the model. The harness scored heuristic_score, which reads two fields off a signal row (strength, occurred_at). Cell 34's production scorer is a BOOSTED_TREE_CLASSIFIER over the nine fv2 features (model/bqml.sql §2). Eight of those nine never reached the reported number. The fallback was beaten, on the same corpus, by a single feature it already computes (strength_sum, AUC 0.7280 unfitted).

  2. The corpus cannot support the gate at all. The fixture converts an identity with probability equal to its latent propensity — a coin flip — so the residual is irreducible Bernoulli noise. Scoring identities by that unobservable propensity gives the oracle ceiling: AUC 0.7277 mean across eight seeds (pooled train estimate 0.7285, 95% CI [0.7134, 0.7435]). The 0.72 autonomy line sits inside the ceiling's confidence interval. No scorer — not the boosted tree, not anything — can reliably clear 0.72 on this fixture. A model that misses it has reported the fixture's signal-to-noise ratio, not its own quality.

A third finding is about the reporting method rather than the model: 0.6884 was one draw of a statistic whose seed-to-seed range is 0.6321–0.7205. It was quoted to four decimals in four documents as though it were a property of the model. On seed 20260812 the unchanged fallback scores 0.7205 and "passes" the same line it "failed" on seed 20260809.

What changed

Measured — eight evaluation seeds, none used in fitting

AUC mean AUC min AUC max Brier mean
baseline (heuristic_score, shipped) 0.6726 0.6321 0.7205 0.2224
model (fv2-logit-1) 0.6967 0.6656 0.7255 0.1845
oracle (unreachable ceiling) 0.7277 0.6690 0.7680 —

On the historically reported seed 20260809 specifically: AUC 0.6884 → 0.7248 (95% CI [0.6535, 0.7903]), Brier 0.2164 → 0.1761. The point estimate clears both 0.70 and 0.72; the interval clears neither, and the report says so in its own output. One seed does not settle it — which is why the panel is the number to cite.

What did NOT work, and is not shipped

Feature engineering beyond fv2 was tested and rejected on measurement. Selected on a train-internal split (fit 1001–1008, validate 1009–1012; evaluation seeds never touched):

feature set validation AUC Brier
oracle ceiling 0.7226 —
fv2 as shipped (10 terms) 0.7077 0.1814
fv2 minus the duplicate column (9 terms) 0.7077 0.1814
fv2 + decay/window/trend features (17 terms) 0.7056 0.1817

Recency-decayed strength at several half-lives, 7/14/30-day signal and strength windows, active-day counts, mean strength and a 7-vs-30-day trend ratio were all tried. None improved on fv2; the enlarged set was slightly worse. fv2 is already close to the corpus's information limit, so shipping an fv3 would have added lockstep burden across scoring.py, bqml.sql and bqml_procedures.sql in exchange for nothing measurable.

A latent defect this surfaced

signals_last_60d is byte-identical to signal_count on 4,982 of 4,982 rows. That is a fixture limitation, not a feature defect: the fixture's observation window is exactly 60 days, so "signals in the last 60d" is necessarily "all signals". In production, where history extends past 60 days, the two differ. The consequence is narrow but real — the backtest cannot exercise that feature at all, and any backtest conclusion about it is vacuous. Fixing it means an observation window longer than the longest feature window; that changes the baseline corpus and so is left as a deliberate follow-up rather than folded into a change whose whole point is comparability.

Tuning that was deliberately NOT done

The fallback's half-life (_OFFLINE_HALF_LIFE_DAYS = 7.0) is mis-set for a 60-day observation window. Measured sensitivity on seed 20260809:

half-life 1d 3d 7d (shipped) 14d 30d 60d 365d
AUC 0.6661 0.6780 0.6884 0.7038 0.7170 0.7201 0.7256

Raising it to 30d would move the headline to ~0.717 with a one-character diff. It is not shipped, for two reasons. Selecting it off the evaluation seed is fitting the constant to the number being reported. And heuristic_score is a serving path (scoring_cell/orchestrator.py:647); shifting its score distribution upward moves identities across the 0.70 enter threshold and changes in-market population on a deployed, IAM-locked service. That is an owner-gated behaviour change with its own hysteresis re-derivation, not a model-quality edit. It is recorded here as a measured option.

Item 6 — what real data would be needed, and what each source buys

The autonomy gate (Brier <= 0.20 AND AUC >= 0.72 AND stable lift over >= 2 purchase cycles) is a real-corpus contract. Nothing on a synthetic fixture can satisfy it, and this change does not claim otherwise. The blocker is not model capacity — it is that there are no forward labels.

backtest/real_data.py already measures and refuses: 605,407 scored rows, 646,065 joining rows, and zero admissible forward labels, because exactly one outcome exists in the entire scoring era. A label is admissible only when score_date < DATE(outcome.occurred_at) <= score_date + horizon; an outcome predating its score is a lookup, not a prediction.

Missing source Register item What it unblocks Expected contribution
Shopify webhooks (orders/create) 6c / arming gate The label stream — forward purchase outcomes at merchant grain Prerequisite, not an improvement. Without it every AUC on real data is undefined, whatever the model.
GA4 export link 6c Session/page depth, funnel-stage signals before purchase Broadens distinct_signal_types and adds pre-purchase depth the current corpus proxies with strength alone.
Klaviyo API key 6c Email engagement (email_open/email_click) as observed rather than fixture-generated Adds a genuinely independent channel; the fv2 aggregates currently collapse all channels into one strength sum.
TTD REDS bucket 6c Ad exposure, and with it the holdout arm Required for lift, the third autonomy criterion, which is not an AUC question at all.
First registered holdout 23 Causal lift over ≥ 2 purchase cycles The only path to the gate's third clause. No activation is legal without a holdout registered before first exposure.

Honest sizing: none of these is quantifiable in AUC terms today, and this report does not estimate one. What is measurable is the shape of the problem — the fixture's ceiling is 0.7277 because the fixture makes conversion a coin flip. Whether real merchant data carries more separable signal than that is an empirical question that the label stream, and only the label stream, can answer.

Recommendations

  1. Stop citing 0.6884 as a model-quality verdict. It is the fallback scorer's score on one draw against a line the corpus cannot support. Cite the panel, or cite nothing.
  2. Re-point the 0.70 harness reference line, or drop it. With an oracle ceiling of 0.7277 the line is not a meaningful target for a synthetic corpus; a ceiling-relative figure ("share of achievable captured") is the honest synthetic metric. Left as an owner decision — this change reports both and enforces neither.
  3. Model quality is no longer the binding constraint on item 17. The binding constraint is the label stream (item 6c) and the holdout (item 23).
  4. Owner decision available: promote fv2-logit-1 onto the offline serving path, which would require re-deriving the hysteresis thresholds against real data. Not done here.

Governance

← All docsView source on GitHub →