Design-vintage (late 2025). The "Active" status column below predates the governance spine and is not a current-state claim. Current design of record for bandit arms, causal bidding and pacing:
docs/design/AD_CONTROL_PLANE_BANDIT_CAUSAL_BIDDING_DESIGN.md(PROPOSED, 2026-09-02); the pacing rule's executable home isservices/service-policy-engine/pacing_veto.py, wired into the policy engine'sevaluate()behindPACING_VETO— a flag that ships off (source literal"false", production env absent), so the veto is dark and nothing on this page is a current-state claim.
Ad Control Plane Strategy: Bandits, Causal Bidding & Pacing Safeguards
Strategic Context
As of late 2025, the MIZ OKI Ad Control Plane is pivoting from simple heuristics to a structured learning framework that balances adaptive online learning (RL/Bandits) with rigorous causal measurement and safety.
This strategy governs the interaction between:
- Cell 3 (KG Brain): Causal inference and strategy selection.
- Cell 11 (Decide ADC): Decision engines (Bandits/RL).
- Cell 15 (Act ADC): Execution and safety guardrails.
The Deployment Checklist (Governing Policy)
All automated bidding implementations must adhere to this 5-step safety progression:
1. Simple Bandit Policies (Phase 1)
- Mechanism: Use Contextual Bandits (e.g., Thompson Sampling, LinUCB) for creative/segment allocation.
- Why: Provides safe, interpretable exploration with theoretical regret bounds.
- Implementation: Current
FunnelOptimizer->DecideADCimplementation.
2. Gated RL with Offline Evaluation (Phase 2)
- Mechanism: Deep RL policies must be evaluated offline ("shadow runs") before controlling live budgets.
- Constraint: RL agents cannot directly control production bids without passing a "Safety Gate" (Actionable by
ReasonADC).
3. Causal Uplift Experiments (Phase 3)
- Mechanism: Run randomized controlled trials (RCTs) or geo-splits to estimate true incremental lift.
- Integration: Feed uplift results into
CausalServiceto calibrate the "Reward" signal for Bandits/RL. - Goal: Optimize for Incremental ROAS, not just attribution-based ROAS.
4. Uplift-to-Bid Translation
- Formula:
Bid = Base_Bid * f(Uplift_Probability) - Requirement: Bids must be shaped by the causal probability of conversion (p(conv|ad) - p(conv|no_ad)).
5. Pacing Safeguards (The "Kill Switch")
- Mechanism: Strict hourly/daily budget pacing controllers.
- Rule: If spend velocity > 200% of target for > 15 mins, trigger Halt.
- Responsibility:
ActADC(Cell 15) must enforce these hard constraints regardless of upstream RL commands.
Architecture Alignment
| Component | Responsibility | Current Status | usage |
|---|---|---|---|
| CausalService (Cell 3) | Estimate Uplift & Confounders | Active (Lagged Regression) | Provide uplift_score to Decide |
| FunnelOptimizer (Cell 3) | Orchestrate SRPVDAL Pipeline | Active | Main entry point for optimization |
| DecideADC (Cell 11) | Bandit/RL Decisioning | Active (Thompson Sampling) | Selects intervention strategy |
| ActADC (Cell 15) | Execution & Pacing | Active | MUST implement strict pacing checks |
| ReasonADC (Cell 7) | Analysis Depth & Safety Checks | Active | Determines if RL is safe to use |
Implementation Priority
- Verify Pacing in ActADC: Ensure
ActADChas hard-coded budget velocity checks. - Refine Uplift Signal: Update
CausalServiceto output an explicituplift_multiplierfor bidding. - Shadow Mode RL: Configure
DecideADCto support a "shadow" mode for testing deeper RL policies without execution.