Ghost-Bid Holdout Testing: Implementation Plan
Document: GHOST_BID_HOLDOUT_IMPLEMENTATION_PLAN.md
Version: 1.0.0
Date: January 30, 2026
Author: MIZ OKI Engineering Team
Status: ๐ IMPLEMENTATION READY
Executive Summary
This document provides a detailed implementation plan for adding ghost-bid holdout testing to the MIZ OKI Autonomous Budget Reallocation MVP. Ghost-bid testing enables true incrementality measurement by randomly withholding ads from a control group and measuring organic conversion rates.
Business Value
| Metric |
Expected Improvement |
| ROAS Accuracy |
Platform-reported โ True incremental (30-50% correction) |
| Budget Efficiency |
2-6% waste reduction |
| Targeting Precision |
+8-15% incremental lift |
| Decision Quality |
Fewer false positives from noisy signals |
1. Architecture Overview
1.1 Current State
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ AUTONOMOUS BUDGET REALLOCATION MVP (V6.13.7) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Layer 1: E-SHKG (Read ROI Deltas) โ
โ Layer 2: ReLU Gating (Pass/Fail Decisions) โ
โ Layer 3: Constrained Optimizer (Generate Plan) โ
โ Layer 4: DSP Execution (Apply Budgets) โ
โ Layer 5: Audit Trail (Firestore Logging) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
1.2 Target State (V6.14.0)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ AUTONOMOUS BUDGET REALLOCATION MVP + GHOST-BID (V6.14.0) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Layer 1: E-SHKG (Read ROI Deltas + Ghost-Bid Deltas) โ
โ Layer 2: Ghost-Bid Test Manager (Create/Track/Analyze) โ โ NEW
โ Layer 3: ReLU Gating (Standard + Ghost-Bid Thresholds) โ โ ENHANCED
โ Layer 4: Incrementality Analyzer (CUPED + Ghost Baseline) โ โ NEW
โ Layer 5: Constrained Optimizer (Live + Holdout Budgets) โ โ ENHANCED
โ Layer 6: DSP Execution (Dual Budget Tracking) โ โ ENHANCED
โ Layer 7: Audit Trail (Test Metadata + Results) โ โ ENHANCED
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
2. Implementation Phases
Phase 1: Data Model Extensions (Week 1)
2.1.1 New Data Classes
GhostBidTest:
@dataclass
class GhostBidTest:
"""Configuration and state for a ghost-bid holdout test."""
test_id: str # Unique identifier
campaign_id: str # Target campaign
name: str # Human-readable name
# Holdout Configuration
holdout_pct: float # Percentage for ghost-bid (default: 2%)
min_ghost_cohort: int # Minimum ghost users (default: 1000)
control_audience_ids: List[str] # Audience segments in ghost group
treatment_audience_ids: List[str] # Audience segments in treatment group
# Timing
start_date: datetime
end_date: Optional[datetime]
duration_days: int # Planned duration
# Status
status: GhostBidTestStatus # DRAFT, ACTIVE, PAUSED, COMPLETED, CANCELLED
# Assignment
assignment_method: str # "hash" or "random"
assignment_seed: str # For reproducible assignments
# Results (populated after analysis)
results: Optional[GhostBidResults]
# Metadata
created_at: datetime
created_by: str
updated_at: datetime
class GhostBidTestStatus(str, Enum):
DRAFT = "draft"
ACTIVE = "active"
PAUSED = "paused"
COMPLETED = "completed"
CANCELLED = "cancelled"
GhostBidResults:
@dataclass
class GhostBidResults:
"""Results from a ghost-bid incrementality test."""
test_id: str
computed_at: datetime
# Sample Sizes
n_treatment: int # Users who saw ads
n_ghost: int # Users who saw ghost (no ad)
# Raw Conversion Rates
treatment_cvr: float # P(convert | ad shown)
ghost_cvr: float # P(convert | no ad) = organic baseline
# Incremental Lift
incremental_lift: float # treatment_cvr - ghost_cvr
incremental_lift_pct: float # (lift / ghost_cvr) * 100
# Confidence Intervals (95%)
lift_ci_lower: float
lift_ci_upper: float
ci_width: float
# Statistical Significance
p_value: float
is_significant: bool # p_value < 0.05
# CUPED Adjustment (if applicable)
cuped_applied: bool
cuped_lift: Optional[float]
cuped_ci_lower: Optional[float]
cuped_ci_upper: Optional[float]
variance_reduction_pct: Optional[float]
# ROI Metrics
incremental_conversions: int
incremental_revenue: float
incremental_roas: float # iROAS = incremental_revenue / treatment_spend
incremental_cpa: float # iCPA = treatment_spend / incremental_conversions
# Recommendation
recommendation: GhostBidRecommendation
recommendation_reason: str
class GhostBidRecommendation(str, Enum):
SCALE_UP = "scale_up" # Significant positive lift โ increase spend
MAINTAIN = "maintain" # Positive but not significant โ continue test
REDUCE = "reduce" # Insignificant lift โ reduce spend
PAUSE = "pause" # Negative lift โ pause campaign
EXTEND_TEST = "extend_test" # Insufficient sample โ continue testing
Enhanced ROIEdge:
@dataclass
class ROIEdge:
"""ROI delta edge with ghost-bid support."""
edge_id: str
audience_id: str
campaign_id: str
# Standard Fields
delta: float # ROI delta (positive = good)
confidence: float # Statistical confidence (0-1)
sample_size: int # Number of observations
# Gating Results
gate_result: GateResult
gated_score: float
# NEW: Ghost-Bid Fields
is_ghost_bid: bool = False # Is this from a ghost-bid test?
test_id: Optional[str] = None # Reference to ghost-bid test
control_audience_id: Optional[str] = None
# NEW: Incremental Metrics
incremental_delta: Optional[float] = None # Lift vs. ghost baseline
incremental_confidence: Optional[float] = None
organic_baseline: Optional[float] = None # Ghost group conversion rate
2.1.2 New Firestore Collections
realloc_ghost_bid_tests:
collection: realloc_ghost_bid_tests
document_id: test_id
fields:
test_id: string
campaign_id: string
name: string
holdout_pct: number
min_ghost_cohort: number
control_audience_ids: array<string>
treatment_audience_ids: array<string>
start_date: timestamp
end_date: timestamp (nullable)
duration_days: number
status: string (enum)
assignment_method: string
assignment_seed: string
results: map (nullable)
created_at: timestamp
created_by: string
updated_at: timestamp
indexes:
- campaign_id, status
- status, created_at DESC
realloc_ghost_bid_assignments:
collection: realloc_ghost_bid_assignments
document_id: auto
fields:
test_id: string
user_id: string
assignment: string ("ghost" | "treatment")
assigned_at: timestamp
hash_value: string
indexes:
- test_id, user_id (unique)
- test_id, assignment
realloc_ghost_bid_outcomes:
collection: realloc_ghost_bid_outcomes
document_id: auto
fields:
test_id: string
user_id: string
assignment: string
converted: boolean
conversion_value: number
conversion_at: timestamp (nullable)
impression_at: timestamp
days_to_conversion: number (nullable)
indexes:
- test_id, assignment
- test_id, converted
Phase 2: Core Components (Week 2)
2.2.1 GhostBidTestManager
class GhostBidTestManager:
"""Manages ghost-bid test lifecycle."""
def __init__(self, firestore_db: Any):
self.db = firestore_db
self.tests_collection = "realloc_ghost_bid_tests"
self.assignments_collection = "realloc_ghost_bid_assignments"
self.outcomes_collection = "realloc_ghost_bid_outcomes"
async def create_test(
self,
campaign_id: str,
name: str,
holdout_pct: float = 0.02,
min_ghost_cohort: int = 1000,
duration_days: int = 7,
control_audience_ids: Optional[List[str]] = None,
created_by: str = "system"
) -> GhostBidTest:
"""Create a new ghost-bid holdout test."""
async def start_test(self, test_id: str) -> GhostBidTest:
"""Start a draft test (DRAFT โ ACTIVE)."""
async def pause_test(self, test_id: str, reason: str) -> GhostBidTest:
"""Pause an active test (ACTIVE โ PAUSED)."""
async def resume_test(self, test_id: str) -> GhostBidTest:
"""Resume a paused test (PAUSED โ ACTIVE)."""
async def complete_test(self, test_id: str) -> GhostBidTest:
"""Complete a test and compute final results."""
async def assign_user(
self,
test_id: str,
user_id: str
) -> str:
"""Assign user to ghost or treatment group (deterministic hash)."""
# Hash-based assignment for reproducibility
hash_input = f"{test_id}:{user_id}"
hash_value = hashlib.sha256(hash_input.encode()).hexdigest()
assignment_int = int(hash_value[:8], 16) % 10000 # 0-9999
test = await self.get_test(test_id)
threshold = int(test.holdout_pct * 10000) # e.g., 2% = 200
assignment = "ghost" if assignment_int < threshold else "treatment"
return assignment
async def record_outcome(
self,
test_id: str,
user_id: str,
converted: bool,
conversion_value: float = 0.0,
impression_at: datetime = None
) -> None:
"""Record conversion outcome for a user."""
async def analyze_test(self, test_id: str, use_cuped: bool = True) -> GhostBidResults:
"""Compute incrementality results for a test."""
2.2.2 IncrementalityAnalyzer
class IncrementalityAnalyzer:
"""Computes incremental lift from ghost-bid tests."""
def __init__(self, cuped_estimator: Optional[CUPEDEstimator] = None):
self.cuped = cuped_estimator or CUPEDEstimator()
async def compute_lift(
self,
treatment_outcomes: List[OutcomeRecord],
ghost_outcomes: List[OutcomeRecord],
use_cuped: bool = True,
pre_period_data: Optional[Dict[str, float]] = None
) -> GhostBidResults:
"""
Compute incremental lift with confidence intervals.
Args:
treatment_outcomes: Conversion data for users who saw ads
ghost_outcomes: Conversion data for ghost-bid users
use_cuped: Apply CUPED variance reduction
pre_period_data: Pre-test conversion rates by user (for CUPED)
Returns:
GhostBidResults with lift, CIs, and recommendation
"""
# Extract conversion rates
n_treatment = len(treatment_outcomes)
n_ghost = len(ghost_outcomes)
treatment_conversions = sum(1 for o in treatment_outcomes if o.converted)
ghost_conversions = sum(1 for o in ghost_outcomes if o.converted)
treatment_cvr = treatment_conversions / n_treatment if n_treatment > 0 else 0
ghost_cvr = ghost_conversions / n_ghost if n_ghost > 0 else 0
# Raw lift
raw_lift = treatment_cvr - ghost_cvr
# Confidence interval (normal approximation)
se_treatment = math.sqrt(treatment_cvr * (1 - treatment_cvr) / n_treatment)
se_ghost = math.sqrt(ghost_cvr * (1 - ghost_cvr) / n_ghost)
se_diff = math.sqrt(se_treatment**2 + se_ghost**2)
ci_lower = raw_lift - 1.96 * se_diff
ci_upper = raw_lift + 1.96 * se_diff
# CUPED adjustment if applicable
cuped_lift = None
cuped_ci_lower = None
cuped_ci_upper = None
vr_pct = None
if use_cuped and pre_period_data:
cuped_result = self.cuped.adjust(
treatment_outcomes, ghost_outcomes, pre_period_data
)
cuped_lift = cuped_result.adjusted_lift
cuped_ci_lower = cuped_result.ci_lower
cuped_ci_upper = cuped_result.ci_upper
vr_pct = cuped_result.variance_reduction_pct
# Statistical significance
z_score = raw_lift / se_diff if se_diff > 0 else 0
p_value = 2 * (1 - stats.norm.cdf(abs(z_score)))
is_significant = p_value < 0.05
# Generate recommendation
recommendation, reason = self._generate_recommendation(
lift=cuped_lift if cuped_lift is not None else raw_lift,
ci_lower=cuped_ci_lower if cuped_ci_lower is not None else ci_lower,
is_significant=is_significant,
n_ghost=n_ghost,
min_cohort=1000
)
return GhostBidResults(
# ... populate all fields
)
def _generate_recommendation(
self,
lift: float,
ci_lower: float,
is_significant: bool,
n_ghost: int,
min_cohort: int
) -> Tuple[GhostBidRecommendation, str]:
"""Generate actionable recommendation."""
if n_ghost < min_cohort:
return (
GhostBidRecommendation.EXTEND_TEST,
f"Ghost cohort ({n_ghost}) below minimum ({min_cohort}). Extend test."
)
if lift > 0.05 and ci_lower > 0 and is_significant:
return (
GhostBidRecommendation.SCALE_UP,
f"Significant positive lift ({lift:.1%}). Recommend scaling budget."
)
if lift > 0 and not is_significant:
return (
GhostBidRecommendation.MAINTAIN,
f"Positive but not significant lift ({lift:.1%}). Continue monitoring."
)
if lift <= 0 and is_significant:
return (
GhostBidRecommendation.PAUSE,
f"Significant negative/zero lift ({lift:.1%}). Recommend pausing campaign."
)
return (
GhostBidRecommendation.REDUCE,
f"Insignificant lift ({lift:.1%}). Consider reducing spend."
)
Phase 3: ReLU Gating Extensions (Week 2-3)
2.3.1 Enhanced Thresholds
@dataclass
class ReallocationThresholds:
"""Thresholds for ReLU gating with ghost-bid support."""
# Standard Thresholds
thresh_uplift: float = 0.05 # 5% minimum uplift
thresh_conf: float = 0.70 # 70% minimum confidence
daily_cap: float = 0.10 # 10% max daily reallocation
max_step: float = 0.20 # 20% max without approval
# NEW: Ghost-Bid Specific Thresholds
ghost_bid_thresh_uplift: float = 0.02 # Lower threshold (2% for holdout)
ghost_bid_thresh_conf: float = 0.60 # Lower confidence (60% vs 70%)
ghost_bid_min_sample: int = 500 # Lower sample requirement
ghost_bid_daily_cap: float = 0.05 # More conservative (5% vs 10%)
# NEW: Incremental ROI Thresholds
min_iroas: float = 1.5 # Minimum incremental ROAS
min_incremental_lift: float = 0.01 # 1% minimum incremental lift
2.3.2 Enhanced Gating Logic
class ReLUGatingOptimizer:
"""ReLU gating with ghost-bid support."""
def gate_edges(
self,
edges: List[ROIEdge],
thresholds: ReallocationThresholds
) -> List[ROIEdge]:
"""Gate edges using appropriate thresholds."""
gated = []
for edge in edges:
# Select thresholds based on edge type
if edge.is_ghost_bid:
thresh_uplift = thresholds.ghost_bid_thresh_uplift
thresh_conf = thresholds.ghost_bid_thresh_conf
min_sample = thresholds.ghost_bid_min_sample
else:
thresh_uplift = thresholds.thresh_uplift
thresh_conf = thresholds.thresh_conf
min_sample = 100 # Default
# Apply ReLU gates
score = self._compute_gated_score(edge, thresh_uplift, thresh_conf, min_sample)
if edge.gate_result == GateResult.PASS:
gated.append(edge)
return gated
def _compute_gated_score(
self,
edge: ROIEdge,
thresh_uplift: float,
thresh_conf: float,
min_sample: int
) -> float:
"""Compute gated score with ghost-bid awareness."""
# Use incremental delta if available (ghost-bid edge)
delta = edge.incremental_delta if edge.incremental_delta is not None else edge.delta
conf = edge.incremental_confidence if edge.incremental_confidence is not None else edge.confidence
# Gate 1: Uplift threshold
if delta < thresh_uplift:
edge.gate_result = GateResult.FAIL_UPLIFT
edge.gated_score = 0.0
return 0.0
# Gate 2: Confidence threshold
if conf < thresh_conf:
edge.gate_result = GateResult.FAIL_CONFIDENCE
edge.gated_score = 0.0
return 0.0
# Gate 3: Sample size
if edge.sample_size < min_sample:
edge.gate_result = GateResult.FAIL_SAMPLE_SIZE
edge.gated_score = 0.0
return 0.0
# All gates pass: compute score
relu_delta = max(0.0, delta)
edge.gated_score = relu_delta * conf * math.log1p(edge.sample_size)
edge.gate_result = GateResult.PASS
return edge.gated_score
| Tool |
Description |
ghost_bid_create_test |
Create new ghost-bid holdout test |
ghost_bid_start_test |
Start a draft test |
ghost_bid_analyze_test |
Compute incrementality results |
ghost_bid_get_recommendation |
Get scaling recommendation |
ghost_bid_assign_user |
Assign user to ghost/treatment |
ghost_bid_list_tests |
List tests for a campaign |
Tool Implementations:
async def ghost_bid_create_test(
campaign_id: str,
name: str,
holdout_pct: float = 0.02,
min_ghost_cohort: int = 1000,
duration_days: int = 7,
firestore_db: Any = None
) -> Dict[str, Any]:
"""
Create a new ghost-bid holdout test for incrementality measurement.
Args:
campaign_id: Target campaign ID
name: Human-readable test name
holdout_pct: Percentage of impressions for ghost-bid (default: 2%)
min_ghost_cohort: Minimum users in ghost group (default: 1000)
duration_days: Planned test duration (default: 7)
Returns:
Test configuration with test_id
"""
manager = GhostBidTestManager(firestore_db)
test = await manager.create_test(
campaign_id=campaign_id,
name=name,
holdout_pct=holdout_pct,
min_ghost_cohort=min_ghost_cohort,
duration_days=duration_days
)
return asdict(test)
async def ghost_bid_analyze_test(
test_id: str,
use_cuped: bool = True,
firestore_db: Any = None
) -> Dict[str, Any]:
"""
Analyze a ghost-bid test and compute incrementality results.
Args:
test_id: Ghost-bid test ID
use_cuped: Apply CUPED variance reduction (default: True)
Returns:
GhostBidResults with lift, CIs, iROAS, and recommendation
"""
manager = GhostBidTestManager(firestore_db)
results = await manager.analyze_test(test_id, use_cuped=use_cuped)
return asdict(results)
async def ghost_bid_get_recommendation(
test_id: str,
firestore_db: Any = None
) -> Dict[str, Any]:
"""
Get scaling recommendation for a ghost-bid test.
Args:
test_id: Ghost-bid test ID
Returns:
Recommendation (SCALE_UP, MAINTAIN, REDUCE, PAUSE, EXTEND_TEST)
"""
manager = GhostBidTestManager(firestore_db)
test = await manager.get_test(test_id)
if test.results is None:
# Compute results if not already done
results = await manager.analyze_test(test_id)
else:
results = test.results
return {
"test_id": test_id,
"recommendation": results.recommendation.value,
"reason": results.recommendation_reason,
"incremental_lift": results.incremental_lift,
"incremental_roas": results.incremental_roas,
"is_significant": results.is_significant,
"ci_lower": results.lift_ci_lower,
"ci_upper": results.lift_ci_upper
}
2.4.2 New API Endpoints (6)
| Endpoint |
Method |
Purpose |
/api/v1/ghost-bid/tests |
GET |
List all tests |
/api/v1/ghost-bid/tests |
POST |
Create new test |
/api/v1/ghost-bid/tests/{id} |
GET |
Get test details |
/api/v1/ghost-bid/tests/{id}/start |
POST |
Start a test |
/api/v1/ghost-bid/tests/{id}/analyze |
POST |
Analyze test results |
/api/v1/ghost-bid/assign |
POST |
Assign user to group |
Phase 5: Integration & Testing (Week 4)
2.5.1 Integration with Reallocation Orchestrator
class ReallocationOrchestrator:
"""Enhanced orchestrator with ghost-bid support."""
async def propose_with_ghost_bid(
self,
campaign_id: str,
test_id: Optional[str] = None,
dry_run: bool = True,
dsp_platform: DSPPlatform = DSPPlatform.DV360
) -> Dict[str, Any]:
"""
Generate reallocation proposal using ghost-bid incrementality data.
Args:
campaign_id: Campaign to reallocate
test_id: Ghost-bid test ID (if active)
dry_run: Preview only (no execution)
dsp_platform: Target DSP
Returns:
Proposal with standard + incremental metrics
"""
# Step 1: Read ROI edges from E-SHKG
standard_edges = await self.eshkg.read_roi_edges(campaign_id)
# Step 2: If ghost-bid test active, read incremental edges
incremental_edges = []
if test_id:
manager = GhostBidTestManager(self.db)
test = await manager.get_test(test_id)
if test.status == GhostBidTestStatus.ACTIVE:
# Compute incremental lift
results = await manager.analyze_test(test_id)
# Convert to ROI edges with incremental data
for audience_id in test.treatment_audience_ids:
incremental_edges.append(ROIEdge(
edge_id=f"ghost_{test_id}_{audience_id}",
audience_id=audience_id,
campaign_id=campaign_id,
delta=results.incremental_lift,
confidence=1.0 - results.p_value, # Convert p-value to confidence
sample_size=results.n_treatment,
is_ghost_bid=True,
test_id=test_id,
incremental_delta=results.incremental_lift,
incremental_confidence=1.0 - results.p_value,
organic_baseline=results.ghost_cvr
))
# Step 3: Combine and gate edges
all_edges = standard_edges + incremental_edges
gated_edges = self.optimizer.gate_edges(all_edges, self.thresholds)
# Step 4: Generate proposal
current_budgets = await self.dsp.get_campaign_budgets(campaign_id)
plan = self.optimizer.propose_reallocation(
campaign_id=campaign_id,
edges=gated_edges,
current_budgets=current_budgets,
total_budget=sum(current_budgets.values())
)
# Step 5: Add incremental metrics to response
return {
"plan": asdict(plan),
"ghost_bid_test": test_id,
"incremental_metrics": {
"incremental_lift": results.incremental_lift if test_id else None,
"incremental_roas": results.incremental_roas if test_id else None,
"recommendation": results.recommendation.value if test_id else None
},
"standard_edges_count": len(standard_edges),
"incremental_edges_count": len(incremental_edges),
"gated_edges_count": len(gated_edges)
}
2.5.2 Test Scenarios
| Scenario |
Expected Behavior |
| Create test with valid params |
Test created with DRAFT status |
| Start test |
Status changes to ACTIVE |
| Assign user (deterministic) |
Same user always gets same assignment |
| Analyze with insufficient data |
Recommendation: EXTEND_TEST |
| Analyze with positive lift |
Recommendation: SCALE_UP or MAINTAIN |
| Analyze with negative lift |
Recommendation: PAUSE |
| Propose with active ghost-bid test |
Incremental edges included in gating |
| Gate ghost-bid edge with lower thresholds |
More lenient gating for incremental data |
3. Rollout Plan
3.1 Canary Stages
| Stage |
Traffic |
Duration |
Criteria to Advance |
| Shadow |
0% |
3 days |
Ghost-bid tests created, no execution |
| Canary |
10% |
7 days |
No errors, valid results |
| Expansion |
25% |
14 days |
iROAS > 1.5, no regressions |
| Production |
70% |
Ongoing |
Continuous measurement |
| Holdout |
10% |
Permanent |
Always run without ghost-bid |
3.2 Feature Flags
| Flag |
Default |
Description |
ENABLE_GHOST_BID_TESTS |
false |
Master switch for ghost-bid testing |
GHOST_BID_DEFAULT_HOLDOUT_PCT |
0.02 |
Default holdout percentage |
GHOST_BID_MIN_COHORT |
1000 |
Minimum ghost cohort size |
GHOST_BID_USE_CUPED |
true |
Apply CUPED variance reduction |
3.3 Monitoring & Alerts
| Metric |
Alert Threshold |
Action |
| Ghost cohort size |
< 500 |
Warn: test may lack power |
| Test duration |
> 30 days |
Warn: consider completing |
| Negative incremental lift |
p < 0.05 |
Alert: recommend pause |
| Assignment imbalance |
> 5% from target |
Error: check randomization |
4. Success Criteria
4.1 Technical Metrics
| Metric |
Target |
| Test creation latency |
< 500ms |
| Assignment latency |
< 50ms |
| Analysis latency |
< 5s |
| Deterministic assignment accuracy |
100% |
4.2 Business Metrics
| Metric |
Target |
| Incremental lift measurement accuracy |
ยฑ1% vs. ground truth |
| CUPED variance reduction |
> 30% |
| Actionable recommendations |
> 80% of completed tests |
| Budget efficiency improvement |
> 5% waste reduction |
5. Dependencies
5.1 Required Components
| Component |
Status |
Notes |
| Autonomous Budget Reallocation MVP |
โ
V6.13.7 |
Base module |
| E-SHKG Client |
โ
V6.11.0 |
ROI edge reading |
| Uplift Pacing Integration |
โ
V5.21.0 |
CUPED estimator |
| Firestore |
โ
Production |
Data persistence |
| DV360 Client |
โ
Mock |
Budget updates |
5.2 Optional Enhancements
| Component |
Status |
Notes |
| Real DV360 API |
๐ก Planned |
Replace mock implementation |
| TTD/Amazon DSP |
๐ก Planned |
Multi-platform support |
| Automated test scheduling |
๐ก Planned |
Recurring tests |
6. Timeline
| Week |
Deliverables |
| Week 1 |
Data models, Firestore collections, basic classes |
| Week 2 |
GhostBidTestManager, IncrementalityAnalyzer |
| Week 3 |
Enhanced ReLU gating, MCP tools, API endpoints |
| Week 4 |
Integration, testing, documentation |
| Week 5 |
Canary rollout (shadow โ 10%) |
| Week 6+ |
Expansion and production rollout |
7. Appendix
7.1 Sample CUPED Calculation
# Pre-period: 14 days before test
# Test period: 7 days
# Step 1: Compute covariance
Y_pre = [0.02, 0.018, 0.021, ...] # Daily CVR before test
Y_post = [0.025, 0.023, 0.024, ...] # Daily CVR during test
cov_Y = np.cov(Y_post, Y_pre)[0, 1]
var_pre = np.var(Y_pre)
# Step 2: Compute adjustment coefficient
theta = cov_Y / var_pre # Typically 0.3-0.7
# Step 3: Adjust post-period
Y_adj = Y_post - theta * (Y_pre - np.mean(Y_pre))
# Step 4: Compare adjusted means
lift_cuped = np.mean(Y_adj_treatment) - np.mean(Y_adj_ghost)
# Variance reduction
vr = 1 - np.var(Y_adj) / np.var(Y_post)
# Typically 30-50% reduction
7.2 Sample Assignment Hash
import hashlib
def assign_user(test_id: str, user_id: str, holdout_pct: float) -> str:
"""Deterministic user assignment."""
hash_input = f"{test_id}:{user_id}"
hash_bytes = hashlib.sha256(hash_input.encode()).digest()
hash_int = int.from_bytes(hash_bytes[:4], byteorder='big')
# Map to 0-9999 range
bucket = hash_int % 10000
threshold = int(holdout_pct * 10000)
return "ghost" if bucket < threshold else "treatment"
# Example:
# assign_user("test_001", "user_12345", 0.02)
# Always returns same result for same inputs
Document prepared by MIZ OKI Engineering Team โข January 30, 2026