ANALYSIS-VINTAGE (July 31, 2026) — retained as a dated record. The §"data stack" line names Neo4j for intent/causal graphs; Neo4j was retired by owner decision 2026-08-09 — the live KG is Firestore-backed and Cell 35's intent graph serves in-memory with a Firestore durability journal (
docs/architecture/CELL_REGISTRY.md). Banner added by the truth-debt sweep, 2026-08-21.
Signal Intelligence Division — Virtuoso Skill-Set Map
Document: SIGNAL_DIVISION_VIRTUOSO_SKILLSETS.md
Version: 1.0
Date: July 31, 2026
Status: Analysis of the built system as of platform v6.45.39 (claim ceiling everywhere: built, pre-benchmark)
Sources analyzed: # MIZ OKI 3.5/signal.html, # MIZ OKI 3.5/demo-signal.html + mizoki_runtime/demo_signal.py, # MIZ OKI 3.5/docs/SIGNAL_INTELLIGENCE_ORACLE_CAPABILITIES.md, Executive Briefing Signal pack, docs/INTENT_API_PLATFORM_BLUEPRINT.md, src/cells/cell33/ (intent-signal-ingest), src/shared/mizoki_intent/, config/intent_taxonomy.yaml, src/cells/google_ads_gaql/ (Cell 29), src/shared/mizoki_media/ (MDES + contracts), src/shared/mizoki_core/ (envelope / objects / control plane), src/shared/miz_oki_source_of_truth.py, src/shared/virtuoso_models/, services/service-validation-orchestrator/, services/service-policy-engine/policy.yaml, miz-oki-adk-agents/boss/ signal tool families, tools/claims_lint.py.
What the Signal Division is
One sentence: "Anticipatory intent with proof of causal lift" — a governed factory that turns raw omnichannel signals (paid search, bidstream, email, CRM, micro-signals) into calibrated intent predictions (ORACLE) and causally-proven budget actions, where nothing acts without passing a ReLU gate, the media validation passport, MDES autonomy bands, and the Decision Control Plane.
Media is the platform's only non-advisory domain (DEL threshold 80, $25k auto-value ceiling, growth-lead approval route per policy.yaml) — a Signal operator carries the most real authority anywhere in MIZ OKI, and correspondingly faces the most gates. A virtuoso operator needs mastery across the ten skill domains below.
The division's operating creed (stated on signal.html itself): predictions inform humans, actions require proof, and the veto is not an opinion — it is arithmetic.
Skill Domain 1 — Causal inference & incrementality science
The division's intellectual core; every claim rests on it.
- Uplift estimation: T-Learner, X-Learner (Künzel 2019), DR-Learner (Kennedy 2020), TMLE, CATE with bootstrap confidence intervals. The four segments and their action mapping: Persuadables (score ≥ 0.8 → BID_UP), Sure Things (≥ 0.5 → MAINTAIN), Lost Causes (≥ 0.2 → BID_DOWN), Sleeping Dogs (else → SUPPRESS/EXCLUDE — treatment harms them).
- Refutation before trust: DoWhy refuters — placebo-treatment, random-common-cause, data-subset. "Proven-incremental" is usable only after refutation passes; the media passport's
causal_refutationcheck fails on absent evidence, never skips. - Model quality metrics: Qini curves, AUUC, calibration (Brier score, ECE, isotonic), discrimination (AUC). Memorized floors: Qini > 0.02, AUUC > 0.55, ATE > 0.005 for cohort activation (
uplift_guardrails.yaml); Brier ≤ 0.20 ∧ AUC ≥ 0.72 ∧ stable lift ≥ 2 purchase cycles for ORACLE model promotion past observe-only — all design targets for promotion gates, with demotion back to observe-only if Brier degrades above 0.20. - Off-policy evaluation: IPS (Horvitz-Thompson), SNIPS, Doubly-Robust (Dudík), with mandatory propensity clipping (max weight 10). Diagnostic instinct: IPS/SNIPS divergence > 0.25 = smart-bidding co-adaptation warning (design target threshold).
- The incrementality evidence ladder as a decision instrument (
mizoki_media.contracts.IncrementalityLadder): ATTRIBUTION_ONLY(0) → OPE_SHADOW(1) → CATE_UPLIFT(2) → EXPERIMENT_CALIBRATED_MMM(3) → GEO_HOLDOUT(4) → LIFT_TEST(5). Budget expansion requires rung ≥ 3 AND interference-adjusted AND iROAS > 1.0 — a high score never substitutes for evidence rung. - iROAS thinking: platform-reported ROAS overstates true incremental return (branded-search iROAS ~0.27×, reported-vs-measured gaps 1.5–3× — published-literature reference points, illustrative scenario in our context). The credit ledger distinguishes CAUSED vs ANTICIPATED conversions, each with a CI.
Skill Domain 2 — Experiment design & measurement
- Holdout, ghost-bid (Ghost Ads, JMR 2017), and geo experiments — knowing which design fits which channel (ghost-bid only inside walled gardens); Wayfair-scale geo splits (~210 DMAs) as the reference pattern.
- Variance reduction: CUPED, switchback designs.
- Holdout hygiene (design targets from
uplift_guardrails.yaml+ blueprint §6): 10% permanent cohort holdout, 5% global holdout, 500-user minimum, 30-day rotation; ghost ads at 2% of impressions / min 1,000 recipients. - MTA + MMM + incrementality triangulation — incrementality is the calibration layer. The
mmm_agreementpassport check requires Meridian-vs-Robyn ROI agreement (complements, report divergence). Hill-function saturation curves and marginal-ROAS equalization drive portfolio allocation. - Leakage & look-ahead-bias discipline: the bitemporal envelope's
available_to_model_atis the only time axis a model may filter on; training examples must use features withavailable_to_model_at <= label_cutoff; the envelope raises by construction ifavailable_to_model_at < observed_at(mizoki_core.envelope, negative-tested in Cell 33's suite).
Skill Domain 3 — Intent modeling / ORACLE
- The four intent stages (Awareness → Consideration → In-market → Purchase-imminent) with hysteresis transitions (enter ≥ 0.70, exit ≤ 0.60 — design targets) and subscription-to-transitions as the operating pattern.
- Two-tower embeddings + sequence models (YouTube RecSys candidate-generation-plus-ranking pattern); BigQuery ML
BOOSTED_TREE_CLASSIFIERfor Phase A, Vertex AI Vector Search for Phase B — with the blueprint cost gate (total ≤ $10K/month, named-human approval for the delta). - Calibrated-probability communication: a well-calibrated "70%" is right ~70% of the time; predictions are probabilities, never promises; every score ships with confidence + explanation path + governance eligibility state.
- Intent graph reasoning (Cell 35
intent-graph, served through Cell 34intent-scoring-api— design, blueprint §5): SHOWED_INTEREST and PRECEDES edges (Neo4j is design-vintage here — retired by owner decision 2026-08-09; the shipped Cell 35 graph is in-memory with Firestore durability); deterministic vs probabilistic household links, each labeled; dark-funnel inference framed explicitly as inference (~17% of the buyer journey is supplier-facing — external-research reference, illustrative scenario here). Probabilistic identities never enter causal math. - Taxonomy governance: the 18-topic × 6-domain
config/intent_taxonomy.yamlis a reviewed commit, never a runtime write; deny-list terms only extendBASE_DENY_TERMS, never shrink; a §8-violating taxonomy makes services refuse to boot — deliberate, don't "fix" it.
Skill Domain 4 — Signal ingestion & data engineering
- Canonical Event Envelope mechanics: deterministic sha256 event_ids (Cell 33 folds
identity_id+occurred_atinto the hashed payload so distinct identities never collide),source_payload_hashcompare-and-set idempotency (inserted / duplicate / updated), provenance stamping, tenant scoping, transactional outbox (DATA-002 —published | pending | dead, never lost). - Connector normalization: JourneyEvent mappers for meta / google_ads / sendgrid / openrtb (
virtuoso_models.transforms),ingest_gateidempotency, and the gateway principle: nothing goes straight from connector to action. - BigQuery modeling: the 8-table
mizoki_intentDDL, partitioning + top-level-column-only clustering (struct subfields are rejected at CREATE TABLE), point-in-time feature builds, counts-only privacy tables (intent_consent_denialsis schema-incapable of holding identities or payloads). - MPP / proxy-open literacy: Apple Mail Privacy Protection detection (ASN / UA / timing signals), click-based engagement validation (opens never promote engagement status), the −0.30 confidence penalty on proxy opens in the demo engine.
- Identity resolution: deterministic-first, probabilistic only where governance permits and always labeled; HMAC-SHA256 pseudonymization that is legally still personal data (GDPR Recital 26); gclid/wbraid/gbraid stitching into CrossPlatformID (90-day window).
Skill Domain 5 — Ad-platform operations
- Google Ads / GAQL as a sensing layer (Cell 29,
src/cells/google_ads_gaql/):SearchStreamextraction,GoogleAdsFieldServicevalidation, the 11-query sensing registry (campaign / keyword / search-term / PMax asset-group / change-status / MCC traversal), GAQL compat-cache semantics (hit / miss / reject / expire with per-account audit ingoogle_ads_gaql_audit), MCC traversal with per-child failure isolation, API version currency (sunset versions hard-fail). - Meta Marketing API: CAPI with 48h
event_iddeduplication, AEM 8-event priority for iOS 14.5+, Custom Audiences / 1% lookalikes, EMQ ≥ 7.0 floor (design target). - Cross-platform hygiene: unified 7d-click / 1d-view attribution windows, Enhanced Conversions (SHA-256-hashed PII), GA4 Measurement Protocol with EU regional endpoints, OpenRTB bidstream, ESP/email mechanics (double opt-in, quiet hours, frequency caps). New offline-conversion work uses Google's Data Manager API (legacy conversion-upload is compatibility-only).
- Smart-bidding levers: conversion value rules, LTV prediction (historical / cohort / ML), Target ROAS optimization — with standing awareness that platform algorithms co-adapt to the signals you feed them.
Skill Domain 6 — The governed decision pathway (SRPVDAL fluency)
The single most important operational skill: knowing exactly what happens at each of the seven stages and which gate can kill an action where.
| Stage | What a virtuoso runs / watches |
|---|---|
| SENSE | Consent-gated ingestion — Cell 33's six-stage pipeline: audio gate → consent gate (fail-closed) → envelope → look-ahead floor → taxonomy validation → idempotent persist → outbox notify |
| REASON / PLAN | All causal reasoning routes through virtuoso_call(Role.DATA_CAUSAL) (currently gemini-3.6-flash via the registry — never hardcode). LLMs draft hypotheses; they never decide |
| VALIDATE | The media validation passport (service-validation-orchestrator): consent_coverage, incrementality_evidence, mmm_agreement, incremental_profit, creative_fatigue, causal_refutation — the full battery always runs; caller-selected subsets are a 422 by design |
| DECIDE | DEL ≥ 80 (passport multiplier: PASS 1.0 / PASS_WITH_CONSTRAINTS 0.75 / FAIL 0.45), MDES autonomy bands 0.55 observe / 0.70 tactical / 0.80 reallocate / 0.90 expand (design targets), the incrementality-ladder check, and an OperatorGate for high-blast-radius actions (budget_expansion, channel_launch/kill, audience_strategy_change, bid_strategy_migration) |
| ACT | Stage 3 recommend-only default; Stage 4 requires all five proofs (evaluation, approval, rollback, monitoring, outcome); a rollback token is minted before every execution; media auto_value_ceiling $25,000 (policy.yaml) |
| LEARN | Predicted-vs-actual OutcomeRecords, EMA weight updates (calibration changes require evidence + a named approver or are refused), drift detection, demotion discipline |
ReLU gate arithmetic every operator should do on a napkin: score = max(0, uplift) × confidence × log(1 + n); floors uplift ≥ 5%, confidence ≥ 0.70, n ≥ 15 conversions; guardrail caps budget ±20% per period, bid ±30%; counterfactual-harm rule: never pause an entity converting at/above account ROAS. (All thresholds: design targets encoded in the guardrail configs.)
Skill Domain 7 — Privacy, consent & compliance engineering
- Fail-closed consent: no affirmative purpose-specific consent → the signal is dropped before persistence and counted (counts only); EU/EEA/UK is deny-by-default and
legitimate_interestis NOT sufficient there; unparseable region strings fall to the strict branch (mizoki_intent.signal). - The three bright lines: no audio-derived signals ever (enforced at validator, taxonomy lint, AND a negative test —
^audiois a 422 that precedes even consent); the sensitive-topic deny-list (health, orientation, religion, politics, ethnicity, immigration, financial distress, union, minors — hard-blocked regardless of confidence); consent-first ingestion. - Regulatory fluency: GDPR Art. 22 (right to human intervention → observe-only default), Art. 17 erasure propagation, CCPA / Meta LDU, TCF 2.2 native strings required in EEA/UK/CH (GPP-wrapped insufficient), no IP-matching in EEA/UK/CH, EU AI Act Art. 14 human-oversight alignment (the MDES error budgets).
- Excluded activation segments: health conditions, financial hardship, political affiliation, religious beliefs, under-13s, vulnerable users.
Skill Domain 8 — Boss-agent MCP tool orchestration
Roughly 100 signal-relevant tools across ~14 families. The virtuoso skill is not knowing each tool — it is knowing which family fires for which situation and how to chain them:
| Situation | Tool family / chain |
|---|---|
| Score & activate an audience | uplift_train_model → uplift_evaluate_model (Qini/AUUC) → uplift_export_cohort (guardrails) → uplift_activate_cohort |
| Reallocate budget | abr_get_roi_edges → abr_propose (ReLU gate) → abr_preview → abr_approve/abr_execute → abr_slo_check / abr_rollback |
| Pace spend on measured lift | uplift_estimate (CUPED/geo/switchback) → uplift_lift_gate_check → uplift_pacing_run (clamp ±20%/period; canary 10 → 20 → 70% with 10% permanent holdout) |
| Cross-channel portfolio | amo_ingest_metrics/mmm/incrementality → amo_optimize → amo_check_drift; bandits via amo_create_experiment / amo_select_arm (Thompson) |
| Creative decay | creative_fatigue_assess (z > 2 CTR drop vs 7-day baseline) → creative_pool_get_replacement (Thompson) → creative_fatigue_rotate (holdback 90/10) |
| Bidding value | vbb_predict_ltv → vbb_add_value_rule → vbb_optimize_target_roas; threshold games via relu_discover_threshold → relu_exploit_threshold (85/15 exploit/explore) |
| Attribution truth | attribution_ingest_event → attribution_record_conversion → attribution_push_weights, plus drift monitoring |
| Journey interventions | journey_compute_risk (DR-learner + PolicyTierGate) → journey_trigger_intervention; stalls via the journey-stall family |
| Closed-loop decisions | clds_canonicalize_event → clds_discover_causal_dag → clds_run_ope → clds_conformal_gate (accept/abstain/escalate) → clds_execute_decision → clds_record_outcome / clds_refute_decision / clds_rollback_decision |
| Graph-native reasoning | gnd_synthesize_evidence / gnd_route_decision / gnd_enforce_policy; per-impression edge scoring with kill switches |
| Intent reads (advisory only) | intent_score_get / intent_cohort_query / intent_transitions_recent / intent_explain / intent_incrementality_report — evidence, never authorization |
The meta-rule: the only activation door for intent is uplift_export_cohort → guardrails → DCP (blueprint §6.6; the measured tool signature is user_ids, uplift_scores, top_k_pct, qini_score, auuc, ate, cohort_name). Intent scores enter decisions as EvidenceBundle items with advisory_only: True.
Skill Domain 9 — GCP platform & production engineering
- Cloud Run fleet ops: IAM-only ingress (unauth
/health→ 403), OIDC minted withformat=fulland the service origin as audience (never a path),ALLOWED_CALLER_SAfail-closed,/healthnot/healthz(GFE-reserved on run.app),latestReady == latestCreatedserving-plane checks,:$COMMIT_SHAimage tags, D8 refuse-to-boot when deployed without a real store. - The data stack: BigQuery + BQML, Firestore (operational evidence), Neo4j (intent/causal graphs), Pub/Sub push with always-ACK poison-message semantics, SSE streaming, Vertex AI Vector Search, Prometheus metrics.
- Virtuoso model registry discipline: never hardcode a model string —
get_model(Role.DATA_CAUSAL); cross-vendor fallback to the global backup;assert_no_legacy_strings()at startup;check_conformance()after any pin or band change.
Skill Domain 10 — Claim discipline & stakeholder communication
The division's brand IS its epistemics, mechanically enforced by tools/claims_lint.py:
- Exactly five labels —
verified result / benchmark result / pilot result / design target / illustrative scenario— required within ±2 lines of any performance number, in any file type (.md,.py,.html,.js, …). - The headline figures (−40% CAC, +35% ROAS, +67% ROI) are design targets; the platform status ceiling is built, pre-benchmark; incrementality harness runs stay
acceptance_eligible: falseuntil a real corpus + named approver exist. - The say / never-say contract: ✅ "calibrated probability this account is in-market" · ❌ "this account will buy" · ❌ "mind-reading" · ❌ any audio-signal claim.
- Executive framing: the CMO/CFO weekly incrementality scorecard, "media spend as governed capital allocation," the caused-vs-anticipated credit ledger, and the five function playbooks (Demand Gen, Lifecycle/CRM, Growth/Performance, Brand, RevOps).
The virtuoso synthesis
What separates a competent operator from a virtuoso is gate arithmetic under pressure: instantly knowing why a specific action was vetoed (the 5% uplift floor? the 15-conversion sample floor? the ±20% budget cap? evidence rung < 3? DEL < 80? a deny-listed segment?) and what evidence would legitimately unlock it — versus what would be a governance violation to even attempt. The system is deliberately built with no bypass parameter anywhere; the virtuoso's power is fluency in earning eligibility (run the right experiment, hit the right evidence rung, mint the rollback, get the named approver) rather than routing around gates.
Known gap (flagged during the dive)
ORACLE's serving cells 34–36 (scoring API at scale, intent graph, causal credit assignment) exist as blueprint + Cell 33 only. Today's virtuoso operates the built half — GAQL sensing (Cell 29), the uplift / attribution / pacing / reallocation tool families, the demo chassis, and Cell 33 consent-gated ingestion — while the intent-scoring half remains checkpointed future work behind the blueprint §9 human-gated rollout phases. Do not deploy cells 33–36 without those human gates.