MIZ OKI 3.0 Cloud Run Resource Budget
Version: 1.0
Last Updated: 2025-11-14
Status: Active
Overview
This document defines the CPU and memory allocation budget for all MIZ OKI 3.0 Cloud Run services. Resource allocations are based on service role, computational requirements, and expected load patterns.
Resource Allocation Philosophy
- Orchestration services (Boss Agent, orchestrators, gateways) receive higher resources to prevent bottlenecks
- Compute-intensive cells (ML inference, optimization, analytics) receive maximum resources
- Standard cells receive baseline resources with autoscaling
- Support services (monitoring, logging) receive moderate resources
Resource Categories
| Category | CPU Request | CPU Limit | Memory Request | Memory Limit | Use Case |
|---|---|---|---|---|---|
| Critical Orchestration | 2 | 4 | 2Gi | 4Gi | Boss Agent, orchestrators, gateways |
| High Compute | 4 | 8 | 8Gi | 16Gi | ML inference, optimization, analytics |
| Standard Cell | 500m | 2 | 512Mi | 2Gi | Most cell services |
| Light Service | 500m | 1 | 512Mi | 1Gi | Simple routing, validation |
Service-Specific Allocations
Orchestration Layer
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| boss-agent | 2 | 4 | 2Gi | 4Gi | Central coordination, 19 agents registered |
| boss-orchestrator | 2 | 4 | 2Gi | 4Gi | SRPVDAL orchestration, prevents bottlenecks |
| orchestration-gateway | 1 | 3 | 1Gi | 2Gi | Gateway routing to MOA/MOE |
| websocket-gateway | 1 | 2 | 1Gi | 2Gi | Real-time WebSocket connections |
| agent-launcher | 500m | 2 | 512Mi | 2Gi | Agent lifecycle management |
SENSE Stage (Cells 1-5)
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell01-discovery | 500m | 2 | 512Mi | 2Gi | Data source enumeration |
| cell02-ingestion | 1 | 2 | 1Gi | 2Gi | Multi-source data extraction |
| cell03-kg-brain | 2 | 4 | 2Gi | 4Gi | E-SHKG operations, graph queries |
| cell04-ontology | 1 | 2 | 1Gi | 2Gi | Semantic reasoning |
| cell05-scheduler | 4 | 8 | 8Gi | 16Gi | Adaptive scheduler, ML/optimization workloads |
REASON Stage (Cells 6-9)
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell06-moe-router | 1 | 2 | 1Gi | 2Gi | Mixture of Experts routing |
| cell07-moa-orchestrator | 1 | 2 | 1Gi | 2Gi | Mixture of Agents orchestration |
| cell08-ml-inference | 2 | 4 | 4Gi | 8Gi | ML model inference |
| cell09-monitoring | 1 | 4 | 2Gi | 16Gi | Observability and metrics |
DECIDE Stage (Cells 10-13)
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell10-validation | 500m | 1 | 512Mi | 1Gi | Data validation |
| cell11-attribution | 1 | 2 | 1Gi | 2Gi | Marketing attribution analysis |
| cell12-planning | 1 | 2 | 1Gi | 2Gi | Strategic planning |
| cell13-simulation | 2 | 4 | 2Gi | 4Gi | What-if scenario simulation |
ACT Stage (Cells 14-18)
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell14-quantum | 2 | 4 | 2Gi | 4Gi | Quantum-inspired optimization |
| cell15-campaign | 1 | 2 | 1Gi | 2Gi | Campaign execution |
| cell16-personalization | 1 | 2 | 1Gi | 2Gi | Personalization engine |
| cell17-creative | 1 | 2 | 1Gi | 2Gi | Creative generation |
| cell18-analytics | 4 | 8 | 8Gi | 16Gi | Advanced analytics, large datasets |
LEARN Stage (Cells 19-25)
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell19-learning | 2 | 4 | 2Gi | 4Gi | Learning engine, model training |
| cell20-pricing | 1 | 2 | 1Gi | 2Gi | Dynamic pricing |
| cell21-router | 500m | 2 | 512Mi | 2Gi | Smart routing |
| cell22-churn | 1 | 2 | 1Gi | 2Gi | Churn prediction |
| cell23-causal | 1 | 2 | 1Gi | 2Gi | Causal discovery |
| cell24-kg-updater | 1 | 2 | 1Gi | 2Gi | Knowledge graph updates |
| cell25-explanation | 1 | 2 | 1Gi | 2Gi | Causal explanation |
Additional Services
| Service | CPU Request | CPU Limit | Memory Request | Memory Limit | Justification |
|---|---|---|---|---|---|
| cell28-neural-processor | 2 | 4 | 2Gi | 4Gi | Neural processing, chat-adjacent |
| cell32-roi-dashboard | 1 | 2 | 1Gi | 2Gi | ROI dashboard and causal analysis |
Autoscaling Configuration
Orchestration Services
- Min Instances: 2 (high availability)
- Max Instances: 8-10
- Target: 80% CPU utilization
- Concurrency: 150-200 requests
High Compute Cells
- Min Instances: 1
- Max Instances: 20
- Target: 80% CPU utilization
- Concurrency: 80
- CPU Throttling: Disabled
Standard Cells
- Min Instances: 1
- Max Instances: 10
- Target: 80% CPU utilization
- Concurrency: 80-100
Cost Optimization Guidelines
- Use CPU-only tiers for services without GPU requirements
- Enable CPU throttling for non-critical services that can tolerate latency
- Set appropriate min instances: 0 for dev, 1 for staging, 1-2 for production
- Monitor actual usage: Review Cloud Monitoring metrics weekly
- Right-size iteratively: Start conservative, scale up based on metrics
Monitoring and Adjustment
Key Metrics to Monitor
-
CPU Utilization - Target: 60-80% average - Alert: >90% for sustained periods
-
Memory Utilization - Target: 60-80% average - Alert: >85% (risk of OOM)
-
Request Latency - Target: p95 <2s for orchestration, <5s for compute - Alert: p95 >5s
-
Instance Count - Monitor actual vs max instances - Adjust max if frequently hitting ceiling
-
Cold Start Frequency - Target: <5% of requests - Consider increasing min instances if higher
Adjustment Process
- Weekly Review: Check resource utilization dashboard
- Monthly Planning: Adjust budgets based on usage trends
- Incident-Driven: Update immediately if bottlenecks occur
- Quarterly Audit: Comprehensive review with cost optimization
Resource Request Guidelines for New Services
When adding new services, use this decision tree:
Is it an orchestration/coordination service?
├─ YES → Category: Critical Orchestration (2 CPU / 4Gi)
└─ NO
├─ Does it perform ML inference or optimization?
│ ├─ YES → Category: High Compute (4-8 CPU / 8-16Gi)
│ └─ NO
│ ├─ Does it process large datasets?
│ │ ├─ YES → Category: Standard Cell (1-2 CPU / 1-2Gi)
│ │ └─ NO → Category: Light Service (500m CPU / 512Mi)
Historical Changes
| Date | Service | Change | Reason |
|---|---|---|---|
| 2025-11-14 | boss-orchestrator | 3 CPU/2Gi → 4 CPU/4Gi | Bottleneck prevention |
| 2025-11-14 | cell28-neural-processor | 3 CPU/2Gi → 4 CPU/4Gi | Chat performance improvement |
| 2025-11-14 | websocket-gateway | - → 2 CPU/2Gi | New service deployment |
| 2025-11-14 | agent-launcher | - → 2 CPU/2Gi | New service deployment |
Budget Summary
Total Resource Budget (Production)
With min instances running: - Total CPU: ~50-60 vCPUs - Total Memory: ~60-80 GiB - Estimated Monthly Cost: $500-800 USD
At max instances (worst case): - Total CPU: ~200-250 vCPUs - Total Memory: ~250-300 GiB - Estimated Monthly Cost: $2,000-3,000 USD
Typical Operating Point: - Total CPU: ~80-100 vCPUs - Total Memory: ~100-120 GiB - Estimated Monthly Cost: $800-1,200 USD
Related Documentation
- Service Discovery Configuration
- Deployment Guide
- Scaling Best Practices (to be created)
- Cost Optimization (to be created)
Maintained by: MIZ OKI Platform Team
Review Frequency: Monthly
Next Review: 2025-12-14