Cloud Run Concurrency & Cold Start Optimization Guide
Version: 1.0.0 Last Updated: January 2026 Scope: MIZ OKI Cloud Run Services Performance Tuning
Overview
This guide provides practical configuration recommendations for optimizing Cloud Run services to minimize latency and cold starts. These settings are critical for the MIZ OKI platform's 32+ microservices architecture.
Key Configuration Parameters
1. Concurrency (Requests Per Instance)
What it is: The maximum number of concurrent requests a single container instance can handle.
| Setting | Range | Default |
|---|---|---|
--concurrency |
1–1000 | 80 (console) |
When to adjust:
| Workload Type | Recommended Concurrency | Rationale |
|---|---|---|
| CPU-heavy (ML inference, image processing) | 1–10 | Each request needs dedicated CPU |
| I/O-bound (API calls, database queries) | 50–200 | Requests spend time waiting |
| Mixed (typical web services) | 20–80 | Balance between CPU and I/O |
CLI:
gcloud run services update SERVICE_NAME --concurrency 10
2. vCPU & Memory Sizing
Available vCPU tiers:
| vCPU | Min Memory | Max Memory | Use Case |
|---|---|---|---|
| 1 | 128 MiB | 4 GiB | Light workloads |
| 2 | 256 MiB | 8 GiB | Standard services |
| 4 | 512 MiB | 16 GiB | Compute-intensive |
| 8 | 1 GiB | 32 GiB | Heavy ML/processing |
CLI:
gcloud run services update SERVICE_NAME --cpu 2 --memory 4Gi
3. Minimum Instances (Keep Warm)
What it is: Number of container instances kept running even with zero traffic.
| Setting | Effect |
|---|---|
min-instances = 0 |
Cold starts on first request (default) |
min-instances = 1 |
One warm instance always ready |
min-instances = 2+ |
Multiple warm instances for high availability |
When to use:
- Latency-sensitive services: Set min-instances >= 1
- Cost-sensitive batch jobs: Keep at 0
- High-traffic services: Scale based on baseline RPS
CLI:
gcloud run services update SERVICE_NAME --min-instances 2
4. Startup CPU Boost
What it is: Temporarily grants extra CPU during container initialization.
Effect: Often halves cold-start time for CPU-bound initialization (dependency loading, JIT compilation, model loading).
How to enable: 1. Google Cloud Console → Cloud Run → Select Service 2. Edit & Deploy New Revision 3. Container tab → Enable Startup CPU boost
Note: This is a console-only setting as of January 2026.
Quick Decision Matrix
| Scenario | Concurrency | vCPU | Min Instances | Startup Boost |
|---|---|---|---|---|
| Boss Agent (mixed I/O + reasoning) | 20 | 2 | 2 | Yes |
| ML Inference Cell (CPU-heavy) | 1–5 | 4 | 1 | Yes |
| API Gateway (I/O-bound) | 100 | 1 | 2 | Yes |
| Batch Processing (async jobs) | 10 | 2 | 0 | No |
| Frontend/UI (static + SSR) | 80 | 1 | 1 | Yes |
MIZ OKI Service Recommendations
Based on the platform architecture:
Tier 1: Always Warm (min-instances >= 2)
| Service | Concurrency | vCPU | Memory | Rationale |
|---|---|---|---|---|
boss-agent-adk |
20 | 2 | 4Gi | Core orchestrator, latency-critical |
miz-oki-command-center-ui |
80 | 1 | 2Gi | User-facing frontend |
api-gateway |
100 | 1 | 2Gi | All traffic entry point |
Tier 2: Single Warm Instance (min-instances = 1)
| Service | Concurrency | vCPU | Memory | Rationale |
|---|---|---|---|---|
cell-03 (KG Brain) |
10 | 2 | 4Gi | Critical data path |
cell-06 (MOE Router) |
20 | 2 | 4Gi | Routing decisions |
ekis |
20 | 1 | 2Gi | Entity enrichment |
Tier 3: Scale to Zero OK (min-instances = 0)
| Service | Concurrency | vCPU | Memory | Rationale |
|---|---|---|---|---|
| Batch processing cells | 5 | 2 | 4Gi | Async, latency tolerant |
| Reporting services | 10 | 1 | 2Gi | Scheduled jobs |
| Dev/staging environments | 10 | 1 | 1Gi | Cost optimization |
Cold Start Reduction Strategies
1. Application-Level Optimizations
# DO: Initialize at module load time (runs once per cold start)
import heavy_library
MODEL = load_model() # Loaded during container init
def handle_request(request):
# Model already loaded
return MODEL.predict(request.data)
# DON'T: Initialize inside request handler
def handle_request_slow(request):
model = load_model() # Loaded on EVERY request
return model.predict(request.data)
2. Dependency Optimization
- Trim unused dependencies: Review
requirements.txt/package.json - Use lazy imports: Only import heavy modules when needed
- Pre-compile Python: Use
.pycfiles in Docker image
3. Docker Image Optimization
# Use slim base images
FROM python:3.11-slim
# Multi-stage builds to reduce image size
FROM builder AS final
COPY --from=builder /app /app
# Copy only what's needed
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY src/ ./src/
4. Connection Pooling
# Cache connections outside request handlers
from google.cloud import firestore
# Global client (reused across requests)
db = firestore.Client()
def handle_request(request):
# Reuses existing connection
doc = db.collection('users').document(request.user_id).get()
Monitoring & Tuning
Key Metrics to Watch
| Metric | Target | Alert Threshold |
|---|---|---|
| Cold start latency | < 2s | > 5s |
| Request latency (p95) | < 500ms | > 2s |
| Instance count | Stable | Oscillating rapidly |
| CPU utilization | 40-70% | > 90% sustained |
Cloud Monitoring Queries
-- Cold start frequency
resource.type="cloud_run_revision"
metric.type="run.googleapis.com/container/startup_latency"
-- Request latency by service
resource.type="cloud_run_revision"
metric.type="run.googleapis.com/request_latencies"
CLI Quick Reference
# View current settings
gcloud run services describe SERVICE_NAME --format="yaml(spec.template.spec)"
# Update concurrency
gcloud run services update SERVICE_NAME --concurrency 20
# Update min instances
gcloud run services update SERVICE_NAME --min-instances 2
# Update max instances
gcloud run services update SERVICE_NAME --max-instances 100
# Update CPU/memory
gcloud run services update SERVICE_NAME --cpu 2 --memory 4Gi
# Combined update
gcloud run services update SERVICE_NAME \
--concurrency 20 \
--min-instances 2 \
--max-instances 50 \
--cpu 2 \
--memory 4Gi
Cost Considerations
| Setting | Cost Impact |
|---|---|
Higher min-instances |
Increases baseline cost (always-on instances) |
Higher cpu |
More expensive per instance-second |
Higher memory |
More expensive per instance-second |
| Startup CPU Boost | Slight increase during cold starts only |
Cost optimization tips:
- Use min-instances = 0 for non-latency-critical services
- Right-size CPU/memory based on actual usage
- Consider Cloud Run jobs for batch workloads (pay only during execution)
References
- Cloud Run Concurrency Documentation
- Configure CPU Limits
- Set Minimum Instances
- Startup CPU Boost Announcement
- Functions Best Practices
Changelog
| Version | Date | Changes |
|---|---|---|
| 1.0.0 | Jan 2026 | Initial guide |