OPT-GOV Phase 2 — creative joint-DML extension + dosage verify
Branch: claude/opt-gov-phase2-dosage-creative · Session: 7pjhmg · Date: 2026-08-22
Authority: OPT-GOV_PRE_APPROVAL_2026-08-21 (standing GATE-1, conditions verified below) + owner in-session triple release 2026-08-22 (approved:merge/commit/deploy) ruling the Prompt-0 collisions as recommended. GATE 2 flag flips remain owner-only; this phase adds no flags.
BUILD 2 — MultimodalCreativeUnbundler → creative_aesthetic.dml_joint (extend-in-place)
Per the Prompt-0 collision ruling, the shipped F1 estimator (dr_effects.py, honest simplified AIPW, per-element) stays untouched; the joint estimator lands beside it in the same package:
- True DML shape: cross-fitted residualization of Y and EVERY component indicator on the covariate buckets (row-level round-robin folds, deterministic; a row's own outcome never enters its nuisance; bucket-less predictions fall back to pooled other-fold means, counted and reported), then a joint impression-weighted least-squares of Y-residuals on all component residuals at once (pure-Python normal equations + Gaussian elimination). Per-component univariate covariance is not offered.
- Identification honesty: components with ~zero residual variance are excluded and reported
unidentified— never a fabricated zero; a residually collinear survivor set reports all-unidentified rather than garbage (exercised in development when a fixture accidentally made the placebo collinear with a real arm). - Collinearity diagnostic: VIF per component over the residualized design; VIF > 10 ⇒
confidence: downgraded_collinear, value always reported. - CIs: deterministic leave-one-bucket-out jackknife over the whole pipeline; a shipped estimate always carries its CI — no-CI components are withheld from findings rows (reported in
diagnostics.withheld_no_ci). - Refutation before surfacing (deterministic forms of the DoWhy refuter families; the module stays pure-Python): a hash-assigned placebo component (refutes when material |θ|>0.005 AND significant) and a data-subset stability check (sign flips above tolerance). A refuted/unstable run returns its status and ships no rows — flagged, never shipped.
- Invariants preserved: findings for REASON/PLAN only (package-wide banned-imports tripwire covers the new module); every row
provisional=True(owner's evidence-backedpilot_validatedflip governs); rows are exactly thef1_element_effectsstore shape (allowlist-enforced); per-component diagnostics travel in the envelope only — persisting them would be an additive DDL (operator-applied), deliberately not done here. - Ontology intake:
ontology/proposals/OCP-20260822-001.mddrafted (steward-level; never self-approved) proposing the closed element vocabulary as a generated SKOS-style scheme, code schema remaining the source of truth.
Tests (tests/test_dml_joint.py, all green in a fresh venv; suite total 78 = 63 existing untouched + 15 new): known effects recovered within tolerance on confounded data (urgency +0.030, people +0.010); correlated null gets no credit jointly while the naive univariate contrast provably does (both directions seeded in one test, with a decay tripwire on the seed); collinearity downgrade; placebo refuter both directions (clean passes; outcome-keyed-to-placebo refutes and ships nothing); subset instability ships nothing; single-bucket data withholds rows for missing CIs; determinism (identical envelopes); store acceptance + provisional lock; balanced-fixture construction is per-cell deterministic so the placebo bit cannot correlate with any arm.
BUILD 1 — ContinuousDosageRLearner: VERIFY-ONLY (collision ruling: skip)
Measured on services/lift-engine/src/core/continuous_dosage.py (landed with GC Activation d58bd79c, 18 tests):
- Cross-fitting ✓ (_cross_fit, group-split folds; METHOD="r_learner_local_marginal_moment_wls_v1").
- CIs ✓ — cluster bootstrap over groups (default 200 refits), with an honest degenerate-refusal (bootstrap_refits_degenerate when <max(20, n/4) refits survive); "every estimate carries bootstrap uncertainty" per its header.
- Never spend-to-max ✓ — refusal statuses (STATUS_INSUFFICIENT_SUPPORT, …_TREATMENT_VARIATION, …_CLUSTERS, …_INVALID_INPUT) instead of extrapolation.
- DoWhy refutation chain: deliberately absent by posture, not by gap — dosage outputs are advisory/non-causal by the activation's settled convention (self-graded CATE minting refused by construction; causal edges only via cells 36→26→27). Routing an advisory summary through DoWhy before PLAN would grade a claim the output deliberately does not make. No code delta.
Gate results (pre-push, fresh venv)
creative_aesthetic 78/78 · governance suite 420 passed · skill_sync.py --audit 0 findings / 14 skills · skills_sync --check + ontology_skills_sync --check both OK · rule-03 V1–V3 greps 0 in-diff · no new flags, no protected paths, no site-visible files · pre-push rebase on current main.