LII backtest model quality — why 0.6884 was the wrong number twice
Date: 2026-08-16
Cell: 34 (intent-scoring-api)
Register item: 17 — "LII offline backtest NO-GO"
Claim label: built, pre-benchmark (unchanged — a synthetic corpus upgrades nothing)
Reproduce:
python3 -m src.cells.cell34.backtest --panel
Summary
Register item 17 recorded the LII offline backtest at AUC 0.6884 against a 0.70 line, and against the 0.72 autonomy gate, and three further documents repeated it. The number is arithmetically correct and was honestly produced. It is nonetheless the wrong number to have been reading, for two independent reasons — and the second one is the important one.
-
It measured the fallback, not the model. The harness scored
heuristic_score, which reads two fields off a signal row (strength,occurred_at). Cell 34's production scorer is aBOOSTED_TREE_CLASSIFIERover the nine fv2 features (model/bqml.sql§2). Eight of those nine never reached the reported number. The fallback was beaten, on the same corpus, by a single feature it already computes (strength_sum, AUC 0.7280 unfitted). -
The corpus cannot support the gate at all. The fixture converts an identity with probability equal to its latent propensity — a coin flip — so the residual is irreducible Bernoulli noise. Scoring identities by that unobservable propensity gives the oracle ceiling: AUC 0.7277 mean across eight seeds (pooled train estimate 0.7285, 95% CI [0.7134, 0.7435]). The 0.72 autonomy line sits inside the ceiling's confidence interval. No scorer — not the boosted tree, not anything — can reliably clear 0.72 on this fixture. A model that misses it has reported the fixture's signal-to-noise ratio, not its own quality.
A third finding is about the reporting method rather than the model: 0.6884 was one draw of a statistic whose seed-to-seed range is 0.6321–0.7205. It was quoted to four decimals in four documents as though it were a property of the model. On seed 20260812 the unchanged fallback scores 0.7205 and "passes" the same line it "failed" on seed 20260809.
What changed
backtest/model.py(new) — the fv2 feature model: an L2-regularised logistic regression fit by IRLS, pure Python, no new dependencies. It is the offline evaluation surrogate for the production tree and explicitly not a serving path;scoring_cell.orchestratorstill servesheuristic_score.backtest/fit.py(new) — fits the coefficients onTRAIN_SEEDSand freezes them intomodel.py. It refuses to run if the training seeds intersect the harness's reporting seeds, so every model number is out-of-sample by construction rather than by good intentions.- Oracle ceiling — the fixture now records the latent propensity that generated each identity, and the harness reports it as an unreachable ceiling plus the share of it captured. A test proves no scorer reads it.
- Bootstrap confidence intervals on every reported AUC, and a
--panelmode across eight evaluation seeds. - Baseline untouched.
heuristic_scoreis unmodified, andresult.aucis still the fallback's, so the 0.6884 in the register remains directly comparable. A test pins it.
Measured — eight evaluation seeds, none used in fitting
| AUC mean | AUC min | AUC max | Brier mean | |
|---|---|---|---|---|
baseline (heuristic_score, shipped) |
0.6726 | 0.6321 | 0.7205 | 0.2224 |
model (fv2-logit-1) |
0.6967 | 0.6656 | 0.7255 | 0.1845 |
| oracle (unreachable ceiling) | 0.7277 | 0.6690 | 0.7680 | — |
- AUC delta: +0.0241 mean, model wins on 7 of 8 seeds.
- Brier: 0.2224 → 0.1845 mean. This is the decisive movement, and it
crosses the autonomy gate's
Brier <= 0.20on 8 of 8 seeds where the fallback crosses it on 0 of 8. The fallback is a ranker; the model emits a calibrated probability. - Share of the achievable ceiling captured: 88% mean (the model recovers most of what the corpus makes recoverable).
On the historically reported seed 20260809 specifically: AUC 0.6884 → 0.7248 (95% CI [0.6535, 0.7903]), Brier 0.2164 → 0.1761. The point estimate clears both 0.70 and 0.72; the interval clears neither, and the report says so in its own output. One seed does not settle it — which is why the panel is the number to cite.
What did NOT work, and is not shipped
Feature engineering beyond fv2 was tested and rejected on measurement. Selected on a train-internal split (fit 1001–1008, validate 1009–1012; evaluation seeds never touched):
| feature set | validation AUC | Brier |
|---|---|---|
| oracle ceiling | 0.7226 | — |
| fv2 as shipped (10 terms) | 0.7077 | 0.1814 |
| fv2 minus the duplicate column (9 terms) | 0.7077 | 0.1814 |
| fv2 + decay/window/trend features (17 terms) | 0.7056 | 0.1817 |
Recency-decayed strength at several half-lives, 7/14/30-day signal and
strength windows, active-day counts, mean strength and a 7-vs-30-day trend
ratio were all tried. None improved on fv2; the enlarged set was slightly
worse. fv2 is already close to the corpus's information limit, so shipping an
fv3 would have added lockstep burden across scoring.py, bqml.sql and
bqml_procedures.sql in exchange for nothing measurable.
A latent defect this surfaced
signals_last_60d is byte-identical to signal_count on 4,982 of 4,982
rows. That is a fixture limitation, not a feature defect: the fixture's
observation window is exactly 60 days, so "signals in the last 60d" is
necessarily "all signals". In production, where history extends past 60 days,
the two differ. The consequence is narrow but real — the backtest cannot
exercise that feature at all, and any backtest conclusion about it is vacuous.
Fixing it means an observation window longer than the longest feature window;
that changes the baseline corpus and so is left as a deliberate follow-up
rather than folded into a change whose whole point is comparability.
Tuning that was deliberately NOT done
The fallback's half-life (_OFFLINE_HALF_LIFE_DAYS = 7.0) is mis-set for a
60-day observation window. Measured sensitivity on seed 20260809:
| half-life | 1d | 3d | 7d (shipped) | 14d | 30d | 60d | 365d |
|---|---|---|---|---|---|---|---|
| AUC | 0.6661 | 0.6780 | 0.6884 | 0.7038 | 0.7170 | 0.7201 | 0.7256 |
Raising it to 30d would move the headline to ~0.717 with a one-character diff.
It is not shipped, for two reasons. Selecting it off the evaluation seed is
fitting the constant to the number being reported. And heuristic_score is a
serving path (scoring_cell/orchestrator.py:647); shifting its score
distribution upward moves identities across the 0.70 enter threshold and
changes in-market population on a deployed, IAM-locked service. That is an
owner-gated behaviour change with its own hysteresis re-derivation, not a
model-quality edit. It is recorded here as a measured option.
Item 6 — what real data would be needed, and what each source buys
The autonomy gate (Brier <= 0.20 AND AUC >= 0.72 AND stable lift over >= 2
purchase cycles) is a real-corpus contract. Nothing on a synthetic fixture
can satisfy it, and this change does not claim otherwise. The blocker is not
model capacity — it is that there are no forward labels.
backtest/real_data.py already measures and refuses: 605,407 scored rows,
646,065 joining rows, and zero admissible forward labels, because exactly one
outcome exists in the entire scoring era. A label is admissible only when
score_date < DATE(outcome.occurred_at) <= score_date + horizon; an outcome
predating its score is a lookup, not a prediction.
| Missing source | Register item | What it unblocks | Expected contribution |
|---|---|---|---|
Shopify webhooks (orders/create) |
6c / arming gate | The label stream — forward purchase outcomes at merchant grain | Prerequisite, not an improvement. Without it every AUC on real data is undefined, whatever the model. |
| GA4 export link | 6c | Session/page depth, funnel-stage signals before purchase | Broadens distinct_signal_types and adds pre-purchase depth the current corpus proxies with strength alone. |
| Klaviyo API key | 6c | Email engagement (email_open/email_click) as observed rather than fixture-generated |
Adds a genuinely independent channel; the fv2 aggregates currently collapse all channels into one strength sum. |
| TTD REDS bucket | 6c | Ad exposure, and with it the holdout arm | Required for lift, the third autonomy criterion, which is not an AUC question at all. |
| First registered holdout | 23 | Causal lift over ≥ 2 purchase cycles | The only path to the gate's third clause. No activation is legal without a holdout registered before first exposure. |
Honest sizing: none of these is quantifiable in AUC terms today, and this report does not estimate one. What is measurable is the shape of the problem — the fixture's ceiling is 0.7277 because the fixture makes conversion a coin flip. Whether real merchant data carries more separable signal than that is an empirical question that the label stream, and only the label stream, can answer.
Recommendations
- Stop citing 0.6884 as a model-quality verdict. It is the fallback scorer's score on one draw against a line the corpus cannot support. Cite the panel, or cite nothing.
- Re-point the 0.70 harness reference line, or drop it. With an oracle ceiling of 0.7277 the line is not a meaningful target for a synthetic corpus; a ceiling-relative figure ("share of achievable captured") is the honest synthetic metric. Left as an owner decision — this change reports both and enforces neither.
- Model quality is no longer the binding constraint on item 17. The binding constraint is the label stream (item 6c) and the holdout (item 23).
- Owner decision available: promote
fv2-logit-1onto the offline serving path, which would require re-deriving the hysteresis thresholds against real data. Not done here.
Governance
- Leakage gate untouched and still load-bearing: 355 planted late-arriving
signals, 0 reach features; removing the gate would report 0.8207
(+0.1323 inflation). Both scorers are gated by the same
compute_featurescall — neither re-derives eligibility. - Probabilistic (household) identities remain excluded from every metric.
- Consent stays fail-closed; the fixture's topics and signal types are still validated through the governed taxonomy.
- The harness still never enforces:
main()exits 0, and a test guards it. - 125 tests pass in
src/cells/cell34/tests/(the 5 modules requiringfastapifail to collect on this host identically at base commit4027cec1— pre-existing, unrelated).