GCP Compute Services Comparison
Document: GCP_SERVICES_COMPARISON.md Last Updated: February 12, 2026 Scope: Cloud Run Services vs Cloud Run Jobs vs Cloud Workflows vs GKE Autopilot
Overview
This document compares the runtime, pricing, and scale tradeoffs between four GCP compute primitives — framed so you can instantly see when each makes sense and where costs and limits apply. Everything here is about choosing the right tool for the job, from fast HTTP APIs to long-running batch and orchestrated flows.
Service Comparison Matrix
| Attribute | Cloud Run Services | Cloud Run Jobs | Cloud Workflows | GKE Autopilot |
|---|---|---|---|---|
| Primary Use Case | Request/response APIs, event-driven endpoints | Batch, background, long-running tasks | Multi-step orchestration across services | Complex distributed apps with strict SLAs |
| Max Timeout | 60 minutes/request | 168 hours (7 days)/task | Up to 1 year (orchestration) | No inherent timeout |
| Scale to Zero | Yes | N/A (task-scoped) | N/A (step-scoped) | No (always-on pods) |
| Autoscaling | Request load + concurrency | Parallel task instances | N/A (not compute) | Pod-level HPA/VPA |
| Billing Model | vCPU-seconds + GiB-seconds (100ms blocks) | Per-instance for full task lifecycle | Per step executed | Always-on pod resources |
| Cold Starts | Yes (mitigated with min instances) | Container startup per task | Negligible (orchestrator) | Minimal (pods scheduled) |
| Concurrency | Up to 1000 per instance | 1 per task instance | Sequential/parallel steps | Pod-level control |
| Ops Overhead | Minimal | Minimal | Minimal | Moderate (K8s concepts) |
| Sidecars | Limited (multi-container preview) | Limited | N/A | Full support |
| Kubernetes Features | None | None | None | Full K8s API |
Cloud Run Services
"Fast, Scalable, Low-Maintenance APIs"
When to Use
- Short-latency request/response workloads (up to 60-minute timeout per request)
- Event-driven APIs that need to scale to zero when idle
- Microservices that autoscale based on request load and concurrency
Key Characteristics
- Concurrency cap configurable up to 1000 per instance — trades off CPU utilization vs per-request latency
- Scale to zero when idle, autoscale up on demand
- Two billing modes:
- Request-based (default): pay only while serving requests — true pay-as-you-go
- Instance-based: non-zero idle CPU to reduce cold starts — always-on charges apply
Pricing
- Billed in vCPU-seconds and GiB-seconds, rounded to 100ms blocks
- Free tier: 180,000 vCPU-seconds/month, 360,000 GiB-seconds/month
- Request-based billing is the most cost-efficient for bursty/intermittent workloads
MIZ OKI Usage
All 32+ microservices (Cells), the Boss Agent (boss-agent-adk), and the Command Center UI run as Cloud Run Services:
| Service | Region | Purpose |
|---|---|---|
boss-agent-adk |
us-central1 | Boss Agent orchestrator |
miz-oki-command-center-ui |
us-central1 | Next.js frontend |
miz-oki-cell{01-32} |
us-central1 | Specialized microservices |
ekis |
us-central1 | Entity Knowledge Intelligence |
Limits to Watch
- 60-minute request timeout (hard limit)
- 32 GiB memory max per instance
- 8 vCPUs max per instance
- 250 max instances per service (adjustable via quota)
Cloud Run Jobs
"Batch and Long-Running Tasks"
When to Use
- Background, batch, or long-running workloads that don't fit request/response patterns
- ETL pipelines, data processing, ML training tasks
- Scheduled tasks (via Cloud Scheduler)
Key Characteristics
- Task timeout up to 168 hours (7 days) — much longer than Cloud Run Services
- Jobs can be parallelized across multiple task instances
- Keeps serverless simplicity while expanding into batch/async workloads
- Billed per-instance for the full lifecycle of the task (no scale-to-zero billing)
Pricing
- Same vCPU-second and GiB-second rates as Cloud Run Services
- Key difference: billed for the entire task duration, not just active request time
- No free tier for Jobs
When to Prefer Over Services
| Scenario | Services | Jobs |
|---|---|---|
| HTTP API endpoint | Yes | No |
| Nightly data pipeline (2h) | Timeout risk | Yes |
| ML model training (6h) | No (60m limit) | Yes |
| Event-driven processing | Yes | No |
| Scheduled batch export | Possible | Better fit |
MIZ OKI Candidates
- Nightly attribution recomputation (
attribution_run_nightly_job) - Daily KG pipeline (
pipeline_kg_run_daily) - Batch scoring (
journey_batch_score) - dbt model runs
- BigQuery DDL migrations
Cloud Workflows
"Orchestration Across Services"
When to Use
- Complex, stateful sequences of tasks across GCP products
- Coordinating Cloud Run Services, Cloud Run Jobs, Pub/Sub, BigQuery, and external APIs
- Long-running multi-step processes with branching logic, retries, and error handling
Key Characteristics
- Not a compute engine — it's a step orchestrator
- Charged per steps executed, not for CPU or instances
- Can run for very long durations (up to a year for overall orchestration)
- Native integration with GCP services (Cloud Run, Pub/Sub, BigQuery, Firestore, etc.)
- Built-in retry, error handling, and conditional branching
Pricing
- $0.01 per 1,000 internal steps (calls to GCP services)
- $0.025 per 1,000 external steps (calls to non-GCP endpoints)
- Free tier: 5,000 steps/month
Orchestration Patterns
Workflow Example: Canary Deployment Gate
┌──────────────────────────────────────────┐
│ 1. Trigger Cloud Run Job: snapshot_a │
│ 2. Wait for deployment (Cloud Build) │
│ 3. Trigger Cloud Run Job: snapshot_b │
│ 4. Call gate-eval-service: /evaluate │
│ 5. Branch: │
│ ├─ promote → Update traffic split │
│ ├─ hold → Wait + re-evaluate │
│ └─ rollback → Revert revision │
└──────────────────────────────────────────┘
MIZ OKI Candidates
- SRPVDAL pipeline orchestration (Sense → Reason → Plan → Verify → Decide → Act → Learn)
- Canary deployment gates (5% → 25% → 100% promotion)
- Cross-platform attribution sync (Google Ads → Meta → GA4)
- Multi-step research agent pipelines
- Budget reallocation approval workflows
GKE Autopilot
"Predictable Kubernetes, Always-On"
When to Use
- Workloads requiring full Kubernetes feature set (sidecars, init containers, DaemonSets)
- Strict p99 latency SLAs where cold starts are unacceptable
- Complex distributed applications with inter-service communication
- Workloads needing GPU/TPU access with fine-grained control
Key Characteristics
- Google manages control plane and node provisioning
- Pricing is always-on — you pay for requested CPU/memory continuously, even during idle
- Full Kubernetes API access with managed infrastructure
- Typically results in higher base costs than Cloud Run but brings full K8s control
Pricing
- Pod-level billing: pay for requested vCPU, memory, and ephemeral storage
- No charge for system pods (managed by Google)
- Higher baseline cost compared to Cloud Run's scale-to-zero model
- Committed use discounts available for sustained workloads
When to Prefer Over Cloud Run
| Requirement | Cloud Run | GKE Autopilot |
|---|---|---|
| Scale to zero | Yes | No |
| Sidecars per pod | Limited | Full support |
| Custom networking (service mesh) | No | Yes (Istio/Anthos) |
| GPU/TPU workloads | Limited | Full support |
| Stateful workloads (StatefulSets) | No | Yes |
| Always-on with strict latency SLAs | Min instances workaround | Native |
| Multi-container pods | Preview | Production |
Decision Framework
Quick Selection Guide
Is it a request/response API?
├─ Yes → Cloud Run Services
└─ No
├─ Is it a batch/background job?
│ ├─ Yes, < 60 minutes → Cloud Run Services (async)
│ └─ Yes, > 60 minutes → Cloud Run Jobs
├─ Is it orchestrating multiple services?
│ └─ Yes → Cloud Workflows
└─ Does it need full Kubernetes features?
└─ Yes → GKE Autopilot
Cost Comparison (Illustrative)
Scenario A: API serving 1M requests/day, 200ms avg latency
| Service | Estimated Monthly Cost |
|---|---|
| Cloud Run Services (request-based) | ~$15-30 |
| GKE Autopilot (always-on pod) | ~$70-150 |
Scenario B: Nightly batch job, 2 hours, 4 vCPU / 8 GiB
| Service | Estimated Monthly Cost |
|---|---|
| Cloud Run Jobs | ~$8-12 |
| GKE Autopilot (CronJob) | ~$70+ (always-on node pool) |
Scenario C: 50-step workflow, 10K executions/month
| Service | Estimated Monthly Cost |
|---|---|
| Cloud Workflows | ~$5 |
| Custom orchestrator on Cloud Run | ~$20-50 (plus maintenance) |
MIZ OKI Architecture Mapping
Current State (All Cloud Run Services)
┌─────────────────────────────────────────────────────────────────┐
│ Cloud Run Services │
├─────────────────────────────────────────────────────────────────┤
│ boss-agent-adk │ miz-oki-command-center-ui │
│ miz-oki-cell{01-32} │ ekis │
│ gate-eval-service │ coding-moa │
│ mizoki-complete │ creative-suite-service │
│ identity-stitcher │ auth-broker │
└─────────────────────────────────────────────────────────────────┘
Recommended Hybrid Architecture
┌─────────────────────────────────────────────────────────────────┐
│ Cloud Run Services (request/response, event-driven) │
│ ├─ boss-agent-adk (API, chat, MCP tools) │
│ ├─ miz-oki-command-center-ui (frontend) │
│ ├─ miz-oki-cell{01-32} (microservices) │
│ ├─ gate-eval-service (deployment gates) │
│ └─ ekis (entity intelligence) │
├─────────────────────────────────────────────────────────────────┤
│ Cloud Run Jobs (batch, long-running) │
│ ├─ Nightly attribution recomputation │
│ ├─ Daily KG pipeline ingestion │
│ ├─ Batch uplift scoring │
│ ├─ dbt model runs │
│ └─ BigQuery DDL migrations │
├─────────────────────────────────────────────────────────────────┤
│ Cloud Workflows (orchestration) │
│ ├─ Canary deployment gates (5% → 25% → 100%) │
│ ├─ Cross-platform attribution sync │
│ ├─ SRPVDAL pipeline orchestration │
│ └─ Budget reallocation approval flows │
└─────────────────────────────────────────────────────────────────┘