MIZ OKI 3.0 Cloud Run Resource Budget

Version: 1.0
Last Updated: 2025-11-14
Status: Active


Overview

This document defines the CPU and memory allocation budget for all MIZ OKI 3.0 Cloud Run services. Resource allocations are based on service role, computational requirements, and expected load patterns.

Resource Allocation Philosophy

  1. Orchestration services (Boss Agent, orchestrators, gateways) receive higher resources to prevent bottlenecks
  2. Compute-intensive cells (ML inference, optimization, analytics) receive maximum resources
  3. Standard cells receive baseline resources with autoscaling
  4. Support services (monitoring, logging) receive moderate resources

Resource Categories

Category CPU Request CPU Limit Memory Request Memory Limit Use Case
Critical Orchestration 2 4 2Gi 4Gi Boss Agent, orchestrators, gateways
High Compute 4 8 8Gi 16Gi ML inference, optimization, analytics
Standard Cell 500m 2 512Mi 2Gi Most cell services
Light Service 500m 1 512Mi 1Gi Simple routing, validation

Service-Specific Allocations

Orchestration Layer

Service CPU Request CPU Limit Memory Request Memory Limit Justification
boss-agent 2 4 2Gi 4Gi Central coordination, 19 agents registered
boss-orchestrator 2 4 2Gi 4Gi SRPVDAL orchestration, prevents bottlenecks
orchestration-gateway 1 3 1Gi 2Gi Gateway routing to MOA/MOE
websocket-gateway 1 2 1Gi 2Gi Real-time WebSocket connections
agent-launcher 500m 2 512Mi 2Gi Agent lifecycle management

SENSE Stage (Cells 1-5)

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell01-discovery 500m 2 512Mi 2Gi Data source enumeration
cell02-ingestion 1 2 1Gi 2Gi Multi-source data extraction
cell03-kg-brain 2 4 2Gi 4Gi E-SHKG operations, graph queries
cell04-ontology 1 2 1Gi 2Gi Semantic reasoning
cell05-scheduler 4 8 8Gi 16Gi Adaptive scheduler, ML/optimization workloads

REASON Stage (Cells 6-9)

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell06-moe-router 1 2 1Gi 2Gi Mixture of Experts routing
cell07-moa-orchestrator 1 2 1Gi 2Gi Mixture of Agents orchestration
cell08-ml-inference 2 4 4Gi 8Gi ML model inference
cell09-monitoring 1 4 2Gi 16Gi Observability and metrics

DECIDE Stage (Cells 10-13)

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell10-validation 500m 1 512Mi 1Gi Data validation
cell11-attribution 1 2 1Gi 2Gi Marketing attribution analysis
cell12-planning 1 2 1Gi 2Gi Strategic planning
cell13-simulation 2 4 2Gi 4Gi What-if scenario simulation

ACT Stage (Cells 14-18)

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell14-quantum 2 4 2Gi 4Gi Quantum-inspired optimization
cell15-campaign 1 2 1Gi 2Gi Campaign execution
cell16-personalization 1 2 1Gi 2Gi Personalization engine
cell17-creative 1 2 1Gi 2Gi Creative generation
cell18-analytics 4 8 8Gi 16Gi Advanced analytics, large datasets

LEARN Stage (Cells 19-25)

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell19-learning 2 4 2Gi 4Gi Learning engine, model training
cell20-pricing 1 2 1Gi 2Gi Dynamic pricing
cell21-router 500m 2 512Mi 2Gi Smart routing
cell22-churn 1 2 1Gi 2Gi Churn prediction
cell23-causal 1 2 1Gi 2Gi Causal discovery
cell24-kg-updater 1 2 1Gi 2Gi Knowledge graph updates
cell25-explanation 1 2 1Gi 2Gi Causal explanation

Additional Services

Service CPU Request CPU Limit Memory Request Memory Limit Justification
cell28-neural-processor 2 4 2Gi 4Gi Neural processing, chat-adjacent
cell32-roi-dashboard 1 2 1Gi 2Gi ROI dashboard and causal analysis

Autoscaling Configuration

Orchestration Services

High Compute Cells

Standard Cells


Cost Optimization Guidelines

  1. Use CPU-only tiers for services without GPU requirements
  2. Enable CPU throttling for non-critical services that can tolerate latency
  3. Set appropriate min instances: 0 for dev, 1 for staging, 1-2 for production
  4. Monitor actual usage: Review Cloud Monitoring metrics weekly
  5. Right-size iteratively: Start conservative, scale up based on metrics

Monitoring and Adjustment

Key Metrics to Monitor

  1. CPU Utilization - Target: 60-80% average - Alert: >90% for sustained periods

  2. Memory Utilization - Target: 60-80% average - Alert: >85% (risk of OOM)

  3. Request Latency - Target: p95 <2s for orchestration, <5s for compute - Alert: p95 >5s

  4. Instance Count - Monitor actual vs max instances - Adjust max if frequently hitting ceiling

  5. Cold Start Frequency - Target: <5% of requests - Consider increasing min instances if higher

Adjustment Process

  1. Weekly Review: Check resource utilization dashboard
  2. Monthly Planning: Adjust budgets based on usage trends
  3. Incident-Driven: Update immediately if bottlenecks occur
  4. Quarterly Audit: Comprehensive review with cost optimization

Resource Request Guidelines for New Services

When adding new services, use this decision tree:

Is it an orchestration/coordination service?
├─ YES → Category: Critical Orchestration (2 CPU / 4Gi)
└─ NO
   ├─ Does it perform ML inference or optimization?
   │  ├─ YES → Category: High Compute (4-8 CPU / 8-16Gi)
   │  └─ NO
   │     ├─ Does it process large datasets?
   │     │  ├─ YES → Category: Standard Cell (1-2 CPU / 1-2Gi)
   │     │  └─ NO → Category: Light Service (500m CPU / 512Mi)

Historical Changes

Date Service Change Reason
2025-11-14 boss-orchestrator 3 CPU/2Gi → 4 CPU/4Gi Bottleneck prevention
2025-11-14 cell28-neural-processor 3 CPU/2Gi → 4 CPU/4Gi Chat performance improvement
2025-11-14 websocket-gateway - → 2 CPU/2Gi New service deployment
2025-11-14 agent-launcher - → 2 CPU/2Gi New service deployment

Budget Summary

Total Resource Budget (Production)

With min instances running: - Total CPU: ~50-60 vCPUs - Total Memory: ~60-80 GiB - Estimated Monthly Cost: $500-800 USD

At max instances (worst case): - Total CPU: ~200-250 vCPUs - Total Memory: ~250-300 GiB - Estimated Monthly Cost: $2,000-3,000 USD

Typical Operating Point: - Total CPU: ~80-100 vCPUs - Total Memory: ~100-120 GiB - Estimated Monthly Cost: $800-1,200 USD



Maintained by: MIZ OKI Platform Team
Review Frequency: Monthly
Next Review: 2025-12-14

← All docsView source on GitHub →