Incident Recovery — Platform Runbook (D6-3 platform half)

claim_label: built, pre-drill. Tooling and procedures are implemented and dry-run-tested (tests/governance/test_incident_scripts.py; the census against a fake gcloud in tests/governance/test_registry_census.py); no live drill has executed them against production (the live-drill half of D6-3 stays gated on EXEC-1's engagement, exactly as docs/runbooks/INCIDENT_RECOVERY_ACTION_CLASSES.md records).

Scope split — read the right runbook:

Every tool that can change anything is dry-run by default, prints exact commands, and refuses --apply where gcloud is absent (cloud sandboxes); registry_census.py makes read calls only, and exits 3 there. Health truth is the control plane, never an HTTP probe: unauth /health 403 = IAM-locked and healthy (OPERATING_SYSTEM 6.3) — do not "recover" a service that is not down.

The six tools (scripts/incident/)

Situation Tool What it does
Is the fleet serving what the registry says? registry_census.py [--iam] [--revisions] [--json] [--only <row>,...] Read-only control-plane census of every production/service-registry.yaml row. Its only calls are gcloud run services describe, plus gcloud run services get-iam-policy with --iam and gcloud run revisions describe with --revisions; it probes no endpoint. A row with a deployed status is ok only when Ready, latestReady == latestCreated, and 100% of traffic is summed on the ready revision; --iam compares the invoker binding with the row's auth:. A row with no recorded service name is a finding (unmapped), never guessed. Exit 1 on any finding, 2 on a usage error, 3 without gcloud, 4 when nothing was checked (no selected row has a deployed status; never a pass). The operator's identity needs run.services.get (plus run.services.getIamPolicy for --iam, run.revisions.get for --revisions). The scheduled Registry Health Check does not run it: that wiring is FR-6's protected-path half, owner-held.
Bad revision serving rollback_revision.py --service <svc> [--to-revision <rev>] Registry-validated per-service rollback: describe → pick previous ready revision → update-traffic pin → verify. --list enumerates the fleet from production/service-registry.yaml (derived service names come from recorded evidence only).
"Is anything armed?" list_kill_switches.py [--json] The global kill-switch / safety-flag inventory with source files and safe defaults — grep-verified against the tree on every run (exit 1 on rot). Live env state still comes from gcloud run services describe.
Runaway/suspect scheduled work pause_schedulers.py [--resume] Pause/resume one-liners for every terraform-declared Cloud Scheduler job, plus the completeness check for jobs created outside terraform.
Drain in flight? drain_status.py Read-only drain/quarantine state commands for gemini-kg-pipeline + the AGENTS 7.7 do-not-touch warning. Run this BEFORE merging anything that touches that service.
Stop all landings deploy_freeze.py The owner-action freeze procedure (disable Auto-Merge + Deploy Router, ledger announcement, optional traffic pins, unfreeze order). Advisory vs. ruleset distinction stated.

Order of operations for a platform incident

  1. Classify. Control-plane read first (rollback_revision.py --service X prints the describe command). 403-unauth is healthy; 503 is real; timeout is a cold start.
  2. Contain. Bad revision → traffic pin to last-good (rollback_revision.py). Suspect ingestion/schedule → pause_schedulers.py. Execution-lane suspicion → the ACTION_CLASSES runbook (kill switch via config redeploy), not this one.
  3. Check the blast surface. list_kill_switches.py for what could be armed; drain_status.py before touching gemini-kg-pipeline.
  4. Freeze if landings would race the fix. deploy_freeze.py procedure; announce on the ledger.
  5. Root-cause into the ledger. python3 scripts/claude_memory.py record with the evidence (failed revision logs age out ~48h — capture FIRST, AGENTS 4.4; if logs are gone, redeploy to capture a fresh crash trace rather than guessing).
  6. Unwind deliberately. Traffic pins make later deploys silent no-ops (GOVERNANCE 6.2) — release pins when fixed; resume schedulers; unfreeze; record the all-clear.

Quarterly drill checklist (dry-run — runnable today, no live risk)

The LIVE drill (executes a real rollback against a real service, and the action-classes drill against a real provider account) is what closes D6-3's ACT half — it runs under an engagement's governance pass (EXEC-1), is owner-scheduled, and its record cites read-back evidence. Until then, statements about this runbook stay at implemented / dry-run-verified.

← All docsView source on GitHub →