Design-vintage (late 2025). The "Active" status column below predates the governance spine and is not a current-state claim. Current design of record for bandit arms, causal bidding and pacing: docs/design/AD_CONTROL_PLANE_BANDIT_CAUSAL_BIDDING_DESIGN.md (PROPOSED, 2026-09-02); the pacing rule's executable home is services/service-policy-engine/pacing_veto.py, wired into the policy engine's evaluate() behind PACING_VETO — a flag that ships off (source literal "false", production env absent), so the veto is dark and nothing on this page is a current-state claim.

Ad Control Plane Strategy: Bandits, Causal Bidding & Pacing Safeguards

Strategic Context

As of late 2025, the MIZ OKI Ad Control Plane is pivoting from simple heuristics to a structured learning framework that balances adaptive online learning (RL/Bandits) with rigorous causal measurement and safety.

This strategy governs the interaction between:

The Deployment Checklist (Governing Policy)

All automated bidding implementations must adhere to this 5-step safety progression:

1. Simple Bandit Policies (Phase 1)

2. Gated RL with Offline Evaluation (Phase 2)

3. Causal Uplift Experiments (Phase 3)

4. Uplift-to-Bid Translation

5. Pacing Safeguards (The "Kill Switch")

Architecture Alignment

Component Responsibility Current Status usage
CausalService (Cell 3) Estimate Uplift & Confounders Active (Lagged Regression) Provide uplift_score to Decide
FunnelOptimizer (Cell 3) Orchestrate SRPVDAL Pipeline Active Main entry point for optimization
DecideADC (Cell 11) Bandit/RL Decisioning Active (Thompson Sampling) Selects intervention strategy
ActADC (Cell 15) Execution & Pacing Active MUST implement strict pacing checks
ReasonADC (Cell 7) Analysis Depth & Safety Checks Active Determines if RL is safe to use

Implementation Priority

  1. Verify Pacing in ActADC: Ensure ActADC has hard-coded budget velocity checks.
  2. Refine Uplift Signal: Update CausalService to output an explicit uplift_multiplier for bidding.
  3. Shadow Mode RL: Configure DecideADC to support a "shadow" mode for testing deeper RL policies without execution.
← All docsView source on GitHub →