MIZ OKI 3.5 Production Convergence Plan
Status: ✅ ADOPTED — validated claim-by-claim against main and integrated July 17, 2026
Branch of record: claude/miz-oki-production-convergence-g2cpci
Validation record: see Appendix A (every P0 claim carries file:line evidence)
Governing rule: all subsequent changes are measured against the release gates, not feature count.
Executive decision
MIZ OKI now enters a production-convergence freeze:
No new product features until the existing platform can be deployed safely, operated reliably, isolated by customer, audited end to end, and onboard a second customer without custom engineering.
The repository is presently suitable for:
- internal demonstrations;
- controlled staging validation;
- recommend-only customer pilots;
- selected live ingestion flows.
It is not yet ready for unrestricted multi-tenant production or autonomous external actions.
The architecture is substantially stronger than the operational implementation. The repository itself identifies missing central cell registration, mixed or partial services, placeholder/dry-run routes, inconsistent SRPVDAL naming, unverified services, and remaining Python 3.9 workloads.
The remediation package (mizoki-remediation/, MIZ-REM-2026-001) has already implemented many of the right primitives — canonical envelopes, validation, policy, approvals, decision control, action authorization, audit/replay, and model registration — but its own documentation says the remaining critical path is deploy, lock down, wire cells, backfill events, and prove domains. Its current end-to-end evidence is primarily eight passing in-memory tests rather than a complete production deployment.
1. Verified critical gaps
P0-1: Production model credentials and fallback are not dependable
The latest live verification (July 13, 2026) found:
- all four Anthropic secrets returned 401;
- the locked CODING_ARCH path therefore had no callable primary;
- the platform-wide Claude fallback reached Anthropic but could not complete because of invalid credentials;
- the
xai-api-keysecret was invalid while a differently named Grok secret (grok-api-key) worked; - OpenAI requests reached the service but failed for insufficient quota;
- most LLM-calling Cloud Run services (85 of 88) lacked the fallback secret mount.
Required fix
- Inventory every model secret and its consumers.
- Disable and delete duplicate or invalid secret names.
- Rotate Anthropic credentials and verify them directly before mounting them anywhere.
- Fund or deliberately disable the OpenAI path until quota is available.
- Replace
xai-api-keywith the verified key and retire the conflicting secret. - Mount each secret only on services that invoke that provider.
- Add startup probes that perform credential validation without generating significant billable traffic.
- Fail closed when a required primary and fallback are both unavailable.
- Rerun one real call for each governed model role and one forced failover test.
- Record the model, provider, latency, token usage, response schema, fallback status, and revision digest as release evidence.
Exit gate: All production model roles either have a verified callable path or are explicitly disabled. No runtime silently falls back to an unverified provider.
P0-2: The deployment system is not safe enough for production promotion
The core fleet workflow (.github/workflows/deploy-all.yml) currently:
- deploys only five named components;
- triggers on a special
.github/fleet-redeployfile or manual dispatch; - authenticates with a stored
GCP_SA_KEY; - does not contain a staging-to-production promotion sequence;
- does not perform canary rollout, automated rollback, or deep post-deployment verification.
The remediation deployment path (mizoki-remediation/ops/deploy_all.sh) builds and deploys :latest images, and its Cloud Build file also publishes only :latest.
Required fix
- Require PRs for every change to
main. - Remove bot paths that merge to
mainwithout required tests. - Protect
mainwith mandatory checks and required review. - Replace stored GCP service-account JSON with GitHub OIDC and Google Workload Identity Federation.
- Restrict the federated identity by immutable repository ID, branch, workflow, and production environment.
- Build each image once and tag it with: Git commit SHA; release version; image digest.
- Stop deploying
latest. - Push images into Artifact Registry rather than the legacy
gcr.iopath. - Generate build provenance, vulnerability results, and an SBOM for every production image.
- Deploy the exact digest to staging.
- Run contract, integration, tenant-isolation, migration, and rollback tests.
- Deploy the same digest to production with zero traffic and a revision tag.
- Run smoke and synthetic tests against the tagged revision.
- Progress traffic through controlled percentages.
- Automatically return traffic to the previous revision if error, latency, or correctness thresholds are breached.
Notes: Workload Identity Federation allows GitHub Actions to exchange its OIDC identity for short-lived Google Cloud credentials instead of retaining a long-lived service-account key. Cloud Run supports tagged revisions, zero-traffic deployments, gradual traffic migration, and immediate rollback to a previous revision. Cloud Build provenance and Artifact Analysis SBOM capabilities depend on storing images in Artifact Registry. Binary Authorization can then prevent untrusted images from being deployed to Cloud Run.
Exit gate: A service can be built, scanned, deployed to staging, promoted by immutable digest, tested at zero traffic, canaried, and rolled back without editing source code or rebuilding the image.
P0-3: Tenant isolation is represented in schemas but not enforced by identity
The canonical envelope contains tenant_id, but service authentication verifies only a Google identity token and optional service-account email allowlist. It does not derive or enforce a tenant.
The Decision Control Plane accepts tenant_id directly from the request body. The approval service similarly accepts a human actor from the request body and returns pending approvals without tenant filtering.
Required fix
Introduce a mandatory CallerContext:
CallerContext
- subject
- service_account
- human_user_id
- tenant_id
- customer_id
- roles
- allowed_domains
- allowed_actions
- authentication_method
- request_id
- trace_id
Then:
- Derive
tenant_idfrom a verified identity mapping, not request JSON. - Reject a body
tenant_idthat differs from the authenticated tenant. - Bind every service call to one tenant unless explicitly operating as an audited platform administrator.
- Derive the approval actor from the verified human identity.
- Implement roles: platform administrator; tenant administrator; operator; approver; auditor; read-only user.
- Require tenant context in every repository/store method.
- Store customer records under tenant-partitioned paths or enforce tenant predicates in every query.
- Scope caches, Pub/Sub attributes, BigQuery rows, graph nodes, audit records, approvals, policies, decisions, and outcomes by tenant.
- Add negative tests that attempt cross-tenant reads, writes, approvals, replay, and actions.
- Prohibit global
list_all()endpoints in customer-facing paths.
Exit gate: A test identity from Customer A cannot observe, infer, approve, replay, modify, or execute anything associated with Customer B — even when it supplies Customer B's IDs directly.
P0-4: Production persistence can silently degrade into process memory
The shared store attempts to initialize Firestore but silently switches to an in-memory store if initialization fails (a warning is logged, but health/readiness is unaffected). That behavior is appropriate for tests but unsafe for production because a nominally healthy service can lose decisions, approvals, authorizations, outcomes, and audit records when the instance restarts.
Required fix
- Add an explicit environment mode:
local;test;staging;production. - Allow memory storage only in
localandtest. - Refuse to start in staging or production when Firestore is unavailable.
- Add
/live,/ready, and/status/deep: -/live: process is alive; -/ready: required storage and configuration are working; -/status/deep: dependency diagnostics for operators. - Add Firestore index and schema deployment to infrastructure as code.
- Add backup, export, restore, and retention policies.
- Add a startup write/read/delete canary against a dedicated health collection.
- Add alarms for fallback attempts, permission failures, latency, contention, and rejected writes.
Exit gate: A production service cannot become ready while using memory storage or while unable to persist and read its required records.
P0-5: The audit chain is vulnerable to concurrent-writer races
The audit implementation separately reads the last hash, separately calculates the next sequence, and then writes the new record. With multiple Cloud Run instances, two writers can observe the same prior record and produce conflicting sequence numbers or parallel chain heads. This is an inference from the current implementation, not a demonstrated incident.
Required fix
- Maintain the chain head in a dedicated Firestore document.
- Create each audit entry and advance the chain head in one Firestore transaction.
- Include: tenant; event time; ingest time; actor identity; service revision; source object hashes; previous chain hash; policy version; contract version.
- Use deterministic IDs or a transactional sequence allocator.
- Export immutable audit copies to BigQuery or Cloud Storage.
- Run periodic chain verification and alert immediately on divergence.
- Add concurrent-writer tests with many simultaneous entries.
- Make audit failure block decisions and actions, not merely log an error.
Firestore transactions provide serializable isolation and are the correct primitive for coordinating concurrent chain-head reads and writes.
Exit gate: Concurrent decision activity preserves one complete, verifiable audit chain, and a forced audit failure prevents authorization or execution.
P0-6: Validation passports can be weakened by the caller
The validation endpoint allows the caller to submit an optional subset of checks. That permits a service to request only favorable validators unless a separate control enforces the complete battery.
Required fix
- Remove arbitrary
checksselection from the production endpoint. - Resolve required checks from: tenant; domain; action class; materiality; policy version; autonomy stage.
- Allow additional checks, never fewer than the required set.
- Include the required-check manifest in the passport seal.
- Verify the passport hash in the Decision Control Plane.
- Cross-check: passport tenant; passport domain; reasoning path; forecast; policy version; expiration; required-check list.
- Reject a passport that was issued before relevant evidence changed.
- Make passports immutable after issuance.
- Add negative tests for omitted checks, changed evidence, mismatched tenants, and tampered hashes.
Exit gate: No caller can choose a smaller validation battery than policy requires, and the DCP independently verifies the passport before deciding.
P0-7: The Stage 4 action path is both nonfunctional and not truly implemented
There are two separate issues.
First, the DCP always issues an authorization marked STAGE_3_RECOMMEND_ONLY. The action runner executes only when neither the registered actuator nor the authorization is Stage 3. Because every issued authorization is Stage 3, the current code can never reach its Stage 4 branch. (Validation note: this is doubly dead — even flipping an actuator to Stage 4 via the registry does not enable execution, because the runner's OR-condition keeps any Stage 3 authorization in the recommend branch, and the DCP never issues anything else.)
Second, even that unreachable Stage 4 branch does not call an external system. It sets executed=True and writes a note. Rollback similarly marks the outcome rolled back without invoking the compensating action.
Required fix
Create a real actuator contract:
ActuatorAdapter
- validate_configuration()
- dry_run(action)
- capture_pre_state()
- validate_preconditions()
- execute(idempotency_key, bounds)
- verify_post_state()
- rollback(pre_state, idempotency_key)
- verify_rollback()
- health()
Then:
- Keep all production actuators at Stage 3 initially.
- Implement one real, low-risk adapter first.
- Have the DCP query the actuator registry and issue Stage 4 only when: tenant policy permits it; the actuator is Stage 4; rollback has been demonstrated; authorization bounds are valid; the action is reversible; the required human or policy approval exists.
- Require both authorization Stage 4 and actuator Stage 4 in the runner.
- Introduce the action lifecycle:
AUTHORIZED → CLAIMED → PRECONDITION_VERIFIED → EXECUTING → EXECUTED
→ POSTCONDITION_VERIFIED → SUCCEEDED
Failure path:
EXECUTING → FAILED → ROLLBACK_PENDING → ROLLED_BACK | ROLLBACK_FAILED
- Claim authorization atomically.
- Use an external-system idempotency key where supported.
- Record pre-state before touching the external platform.
- Verify the actual post-state instead of trusting an API response.
- Automatically roll back on failed postconditions.
- Add a global kill switch and per-tenant/per-actuator kill switches.
- Never report
executed=Truewithout external proof.
Exit gate: A sandbox or narrowly bounded production action can be executed, verified, rolled back, and replayed from the audit record. A simulated action is labeled as simulated and never recorded as executed.
P0-8: Decision signing and approval attribution need stronger trust boundaries
The DCP and action runner both default to a hardcoded development signing key when the environment variable is absent. The deployment script mounts the same signing secret broadly across all ten remediation services rather than only the issuer and verifier. (Validation note: ops/deploy_all.sh does generate and mount a real Secret Manager key in the deployed posture; the hardcoded default applies whenever the env var is absent — local runs, misconfigured deploys, or any other execution path — and the broad mount grants signing authority to services that should only verify.)
Required fix
- Remove every default signing key.
- Refuse startup when signing configuration is absent.
- Prefer asymmetric Cloud KMS signing: DCP receives permission to sign; action runner receives permission to verify or uses the public key; no other service receives signing authority.
- Include in the signed authorization: tenant; decision; action; target resource; precondition version; bounds; nonce; issued time; expiration; policy version; approver identity; permitted actuator.
- Bind human approval identity to verified login claims.
- Require separation of duties for high-materiality actions.
- Log every signing, verification, refusal, and key-version transition.
- Test expired, replayed, cross-tenant, altered, wrong-key, and revoked authorizations.
Exit gate: Only the DCP can create a valid authorization; a service or customer cannot impersonate an approver or alter authorization fields.
P0-9: Canonical events remain fragmented
The active README acknowledges that canonical event schemas, KG mapping tests, and a central Intelligence Cell registry remain incomplete. The codebase currently contains multiple event representations and persistence paths (virtuoso JourneyEvent, mizoki_core envelope, mizoki_contracts envelope, connector-gateway CanonicalBusinessEvent). The remediation documentation also says historical BigQuery and Neo4j data still need enveloping and backfilling.
Required fix
Do not force every domain payload into one flat schema. Adopt a layered contract:
CanonicalEventEnvelope
├── universal identity, tenant, time, source, provenance and hash
├── payload_schema_id
├── payload_schema_version
└── payload: versioned domain event
Domain payload
├── marketing JourneyEvent
├── finance event
├── CRE event
├── legal/policy event
└── operational event
Projections
├── Firestore operational record
├── BigQuery analytical table
└── Neo4j KG nodes and edges
Implementation sequence:
- Write an architecture decision record naming the canonical envelope.
- Establish a schema registry with compatibility rules.
- Assign every current event producer to: canonical; adapter required; duplicate; archive.
- Create contract tests and golden vectors.
- Add adapters from existing formats.
- Dual-write the old and new projections.
- Reconcile counts, hashes, tenant IDs, and timestamps.
- Backfill historical events conservatively using first known ingest time for unknown model availability.
- Rebuild KG projections from canonical events.
- Cut readers over to the canonical path.
- Stop legacy writes.
- Retain raw evidence and migration manifests for replay.
Exit gate: One source event produces one stable canonical ID that can be traced consistently through ingestion, BigQuery, the KG, reasoning, validation, decision, action, outcome, and learning.
P0-10: Ingestion lacks a durable publish/retry boundary
Canonical ingestion treats Pub/Sub as optional, persists the event, and initiates a publish without awaiting or recording the publish result. It also retrieves point-in-time evidence by iterating over all events.
Required fix
- Store the canonical event and an outbox item atomically.
- Run an outbox publisher that: publishes with the canonical event ID; waits for acknowledgement; records attempt count; retries safely; routes permanent failures to a dead-letter queue.
- Add an idempotent consumer contract.
- Add schema version and tenant attributes to messages.
- Alert on outbox age, retry count, DLQ depth, and processing lag.
- Replace full collection scans with indexed tenant/domain/time queries.
- Add replay tooling by tenant, event family, time range, and schema version.
- Mark
/readyunhealthy if a mandatory downstream dependency is unavailable beyond the allowed degradation policy.
Exit gate: Restarting or failing any ingestion component does not lose an event, duplicate an effect, or mix tenants.
2. Ordered production implementation program
Wave 0 — Freeze, inventory, and establish one baseline
Actions
- Change product status from unrestricted "Production" to Production Candidate / Controlled Pilot until all release gates pass.
- Freeze feature development.
- Create a single production-readiness ledger covering every deployed service.
- Inventory all Cloud Run services, jobs, topics, subscriptions, datasets, Firestore collections, graph databases, secrets, schedulers, service accounts, and public endpoints.
- Classify every service: production-required; pilot-required; internal tooling; duplicate; archive; unknown.
- Record owner, source path, deployment workflow, runtime, service identity, tenant model, dependencies, health endpoints, SLO, and rollback procedure.
- Disable or quarantine orphaned services and unknown public endpoints.
- Establish one golden customer use case for production proof.
Deliverable: production/service-registry.yaml becomes the operational source of truth and generates deployment matrices, IAM rules, monitoring, and verification.
Gate: No unknown service, ownerless service, unexplained public ingress, or undocumented production dependency remains.
Wave 1 — Protect the repository and rebuild CI/CD
Actions
- Enforce protected
main. - Require: review; passing tests; no unresolved review threads; no direct pushes; no unverified bot merges.
- Add path-aware CI generated from the service registry.
- Run: formatting and lint; type checks; unit tests; contract/schema tests; tenant-isolation tests; claims lint; secret scanning; dependency scanning; container build; manifest validation.
- Replace
GCP_SA_KEYwith Workload Identity Federation. - Replace
latestwith immutable image digests. - Move all production images to Artifact Registry.
- Generate provenance and SBOMs.
- Add vulnerability severity gates.
- Enable Binary Authorization after the image pipeline is stable.
- Add staging and production GitHub environments with separate approvals.
- Generate a signed release-evidence manifest.
Gate: Nothing can reach production unless it passed the exact tests associated with the exact image digest being deployed.
Wave 2 — Repair credentials and runtime dependencies
Actions
- Complete the model credential work described under P0-1.
- Migrate remaining deprecated model SDK paths.
- Remove dead PaLM calls and stale model strings.
- Route remaining high-value direct calls through the governed dispatcher (
virtuoso_call). - Remove prohibited generation parameters as those paths migrate.
- Upgrade Python 3.9 images (cell26, cell23) to Python 3.11 or newer supported bases.
- Lock dependencies and test dependency resolution before Cloud Build.
- Add startup model-registry and credential guards to every LLM-calling service.
- Ensure a provider failure is visible through metrics and does not silently change the decision's evidence class.
Gate: Every active LLM path is registry-controlled, startup-validated, observable, and covered by a real provider test.
Wave 3 — Enforce service identity, human identity, and tenant isolation
Actions
- Implement
CallerContext. - Create one service account per production service or tightly coupled trust group.
- Generate the IAM call graph from the service registry.
- Replace the blanket auth-hardening script with exposure classes: public website; authenticated customer gateway; internal service; scheduled job; administrative endpoint.
- Bind tenant and roles at the gateway.
- Propagate signed tenant context internally.
- Prevent downstream services from accepting an arbitrary tenant override.
- Split human approvals from service-to-service authentication.
- Make every data-access method tenant-scoped.
- Run a cross-tenant penetration test.
Gate: All production calls have authenticated service and tenant identities, and all high-risk actions have authenticated human attribution.
Wave 4 — Deploy and harden the governance control plane
Actions
- Deploy the ten remediation services to staging.
- Eliminate in-memory fallback in staging and production.
- Deploy dedicated service identities.
- Configure inter-service URLs and audiences.
- Implement transactional audit writes.
- Implement KMS-backed authorization signing.
- Fix validation battery enforcement.
- Bind passports, policies, decisions, approvals, authorizations, outcomes, and learning records by tenant and object lineage.
- Version tenant policies as immutable artifacts.
- Restrict policy publishing and reload operations to authorized roles.
- Add dependency-aware readiness endpoints.
- Run the full governed pathway against real Firestore, not the memory store.
Gate: The real GCP pathway — not an in-memory substitute — proves envelope → passport → policy → decision → approval → authorization → recommendation → audit replay.
Wave 5 — Converge canonical data and KG projections
Actions
- Approve the canonical event architecture decision.
- Build the schema and cell registries.
- Add an outbox and DLQ.
- Add canonical mapping tests per connector.
- Implement dual-write migration.
- Reconcile existing Firestore, BigQuery, and graph representations.
- Backfill historical data.
- Add graph integrity checks: orphan nodes; duplicate entities; missing provenance; stale evidence; cross-tenant edges; schema drift.
- Prove deterministic replay from canonical events.
- Remove duplicate active stores after cutover.
Gate: The same decision can be reconstructed from immutable source evidence and produces an identical governed explanation on replay.
Wave 6 — Implement one real, reversible action path
Actions
- Choose one low-risk actuator.
- Implement the adapter contract.
- Add customer-specific bounds and policy.
- Run dry-run comparisons against the current external state.
- Capture pre-state.
- Execute in a sandbox or narrowly bounded production scope.
- Verify post-state.
- Run an intentional rollback drill.
- Verify rollback state.
- Add action and rollback alarms.
- Keep every other actuator recommend-only.
Gate: One real actuator is proven end to end. No other actuator receives bounded autonomy merely because its endpoint exists.
Wave 7 — Establish SRE operations
Actions
- Add OpenTelemetry trace propagation across the full SRPVDAL path.
- Standardize structured log fields: tenant; trace; event; decision; phase; service; revision; policy; model; provider; action.
- Build dashboards for: availability and latency; event freshness and throughput; schema rejection; queue and outbox age; model errors and fallbacks; validation failures; approvals; action and rollback state; tenant-level data quality.
- Create alert sinks for email, Slack or Teams, and incident escalation.
- Write runbooks for: bad deployment; invalid secret; provider outage; Firestore failure; Pub/Sub backlog; corrupted event; cross-tenant risk; failed action; failed rollback.
- Conduct failure-injection tests.
- Conduct load and soak tests.
- Restore data from backups.
- Roll back a Cloud Run release as a drill.
Gate: An operator can detect, diagnose, contain, roll back, and document a failure without relying on the original developer.
Wave 8 — Build the customer onboarding factory
The onboarding system should be operational automation and documentation — not a large new product module.
Customer manifest — each customer receives a version-controlled manifest:
customer:
customer_id:
tenant_id:
legal_name:
environments:
data_region:
support_tier:
identity:
identity_provider:
administrators:
approvers:
operators:
auditors:
connectors:
- type:
account:
scopes:
credential_secret:
sync_frequency:
historical_start:
data_owner:
objectives:
- metric:
target:
window:
owner:
guardrails:
- policy:
threshold:
approval_route:
autonomy:
default: observe
permitted_actions: []
prohibited_actions: []
success:
first_event:
data_quality:
first_recommendation:
acceptance_rate:
business_outcome:
Onboarding sequence
- Commercial and use-case qualification — one narrow business outcome; named executive owner; named operational owner; measurable acceptance criteria.
- Security and data review — data inventory; PII classification; retention; geographic constraints; access scopes; incident contacts.
- Tenant provisioning — tenant ID; identity mapping; service accounts; secrets; datasets; graph partition; policies; dashboards.
- Connector certification — authentication; API version; permissions; rate limits; pagination; freshness; deduplication; schema mapping.
- Historical backfill — immutable raw evidence; canonical conversion; reconciliation report; data-quality sign-off.
- Objective and guardrail workshop — business objectives; decision thresholds; prohibited actions; approvers; rollback conditions.
- Shadow mode — ingest and reason; no external action; compare recommendations with existing decisions; calibrate confidence.
- Recommend-only pilot — operator receives recommendations; every acceptance or rejection is recorded; predicted and actual outcomes are compared.
- Limited execution — only after rollback proof; only for one action class; hard financial and scope bounds; immediate kill switch.
- Customer acceptance — data quality accepted; security accepted; audit replay demonstrated; operators trained; incident contacts tested; rollback demonstrated.
- Production promotion — signed go/no-go record; release digest; policy version; customer manifest version; rollback target; support coverage.
- Post-launch operations — daily operational review during initial production; weekly value and data-quality review; monthly policy and access review; documented expansion criteria.
Onboarding gate: The second customer must be onboarded from the same templates and automation without editing core application logic.
3. Production release gates
MIZ OKI reaches production readiness only when all of the following are true:
| Gate | Required proof |
|---|---|
| Repository safety | Protected main, mandatory CI, no unchecked bot merges |
| Build integrity | Immutable digest, provenance, SBOM, vulnerability pass |
| Deployment | Staging proof, zero-traffic test, canary, rollback |
| Credentials | All active primary/fallback paths verified |
| Service identity | Least-privilege IAM and intended ingress only |
| Tenant isolation | Cross-tenant negative suite passes |
| Persistence | No production memory fallback |
| Audit | Transactional, tamper-evident, concurrent-write proof |
| Canonical data | One envelope and versioned payload contracts |
| Validation | Required batteries cannot be bypassed |
| Decision control | Object lineage and tenant binding verified |
| Approval | Human identity derived from authentication |
| Action | Real adapter, idempotency, pre/post-state proof |
| Rollback | Compensating action executed and verified |
| Observability | Metrics, traces, dashboards, alerts and runbooks |
| Reliability | Failure injection, load, soak and restore proof |
| Customer | UAT, operator training, acceptance and support plan |
| Repeatability | Second tenant provisioned without core-code changes |
4. First engineering merge sequence
Implement in this order:
| Order | Ticket | Result |
|---|---|---|
| 1 | PROD-001 | Freeze features and classify product as controlled pilot |
| 2 | OPS-001 | Complete service/resource inventory and registry |
| 3 | CI-001 | Protect main and require tests before merge |
| 4 | CI-002 | Replace GCP_SA_KEY with Workload Identity Federation |
| 5 | CI-003 | Replace latest with Artifact Registry digest deployment |
| 6 | SEC-001 | Rotate provider credentials and clean Secret Manager |
| 7 | STORE-001 | Remove production memory fallback |
| 8 | AUDIT-001 | Make audit append transactional |
| 9 | TENANT-001 | Add authenticated CallerContext and tenant enforcement |
| 10 | APPROVAL-001 | Bind approver identity to authentication |
| 11 | VALIDATE-001 | Remove caller-selected validation subsets |
| 12 | DCP-001 | Verify passport hash, tenant, domain and lineage |
| 13 | ACTION-001 | Correct Stage 4 authorization logic |
| 14 | ACTION-002 | Implement the first real adapter and rollback |
| 15 | DATA-001 | Approve canonical event ADR and schema registry |
| 16 | DATA-002 | Implement outbox, retry and DLQ |
| 17 | OBS-001 | Add end-to-end tracing and production dashboards |
| 18 | ONBOARD-001 | Create customer manifest and provisioning workflow |
| 19 | PILOT-001 | Run the golden customer flow in shadow mode |
| 20 | RELEASE-001 | Complete the production go/no-go evidence package |
Final production principle
MIZ OKI should not be declared production-ready because all the components exist. It should be declared production-ready when:
One canonical event can enter through an authenticated tenant boundary, survive dependency failures, update the graph, produce a validated decision, obtain attributable approval, execute a bounded reversible action, verify the real outcome, learn from it, and be replayed from an intact audit trail — then the same process can be provisioned for another customer without changing core code.
The first merge should be PROD-001 plus OPS-001; every subsequent change should be measured against the release gates above rather than against feature count.
Appendix A — Validation record
Every P0 claim was independently verified against the repository on July 17, 2026 (branch claude/miz-oki-production-convergence-g2cpci, based on main @ 3aca2b7). All 10 P0 claims validated.
| Claim | Evidence | Verdict |
|---|---|---|
| P0-1 — model credentials undependable | CLAUDE.md v6.45.14 live verification: all four Anthropic secrets 401; xai-api-key invalid (grok-api-key works); OpenAI insufficient_quota; only 3 of 88 Cloud Run services mount ANTHROPIC_API_KEY. Re-confirmed open in v6.45.18 (July 17). |
✅ Confirmed |
| P0-2 — deploy system unsafe for promotion | .github/workflows/deploy-all.yml:3-14 (fleet-redeploy trigger, 5 services only), :34 et al. (credentials_json: ${{ secrets.GCP_SA_KEY }}); no staging/canary/rollback stages anywhere in the workflow. mizoki-remediation/ops/deploy_all.sh:24-25 deploys gcr.io/$PROJECT/$NAME:latest; ops/cloudbuild.yaml tags and publishes only :latest. |
✅ Confirmed |
| P0-3 — tenant not enforced by identity | mizoki-remediation/contracts/mizoki_contracts/auth.py:28-52 verifies Google OIDC token + optional SA email allowlist only — no tenant derivation. DCP service-decision-control-plane/main.py:46 takes tenant_id from request body. Approval service-approval-routing/main.py:32 takes actor as a plain body string; :84-86 /api/v1/pending returns all pending approvals with no tenant filter. |
✅ Confirmed |
| P0-4 — silent memory fallback | contracts/mizoki_contracts/store.py:32-38 — Firestore init failure logs a warning and falls back to self._mem; no readiness impact. (Precision note: a warning is logged; "silent" is accurate with respect to health/readiness and callers.) |
✅ Confirmed |
| P0-5 — audit chain race | store.py:68-89 — _last_chain_hash() and _next_seq() are two separate non-transactional Firestore reads followed by a plain write in audit() (:91-111). Correctly framed in the plan as an inference (race window), not an observed incident. |
✅ Confirmed |
| P0-6 — caller-weakened validation | service-validation-orchestrator/main.py:120 (checks: Optional[List[str]] — "subset; default = full domain battery") and :135 (names = req.checks or sorted(battery)). |
✅ Confirmed |
| P0-7 — Stage 4 dead + simulated | DCP main.py:85 hardcodes stage=AutonomyStage.STAGE_3_RECOMMEND_ONLY on every authorization. Runner service-action-runner/main.py:135 — if stage == STAGE_3 or auth.stage == STAGE_3: makes the Stage 4 branch unreachable (auth.stage is always Stage 3). The branch itself (:139-143) sets executed=True with a note and the comment "Real actuator adapters plug in here" — no external call. /rollback (:160-175) sets rolled_back=True and appends a note without invoking the compensating action. |
✅ Confirmed (understated: doubly dead) |
| P0-8 — signing trust boundaries | Hardcoded default "dev-only-key-rotate-in-kms" in DCP main.py:36 and runner main.py:29. ops/deploy_all.sh:16-29 creates a random Secret Manager key but mounts DCP_SIGNING_KEY=dcp-signing-key:latest on all ten services in the loop — signing authority granted to services that should only verify (or need neither). |
✅ Confirmed |
| P0-9 — canonical events fragmented | Multiple active event contracts coexist: src/shared/virtuoso_models/journey_event.py (JourneyEvent), src/shared/mizoki_core/envelope.py (CanonicalEventEnvelope), mizoki-remediation/contracts/mizoki_contracts/envelope.py (second CanonicalEventEnvelope), miz-oki-adk-agents/boss/connector_gateway/ (CanonicalBusinessEvent). Remediation README lists backfill of BigQuery/Neo4j as remaining. Repo docs record mizoki_contracts vs mizoki_core reconciliation as explicit future work (CLAUDE.md v6.45.17). |
✅ Confirmed |
| P0-10 — no durable publish boundary | service-canonical-ingestion/main.py:28-34 — Pub/Sub client optional (warning, still "healthy"); :82-84 — _publisher.publish(...) future never awaited or recorded (fire-and-forget, no outbox/retry/DLQ); :88-97 — point-in-time query iterates STORE.list_all("events") (full collection scan). |
✅ Confirmed |
| "eight passing in-memory tests" | mizoki-remediation/tests/test_end_to_end.py — exactly 8 test functions; line 14 forces MIZOKI_STORE=memory. |
✅ Confirmed |
Additional supporting facts verified:
- Six-domain event/decision state is in-memory per revision in production today (
firestore_persistence: falseper livesixdomain_status, July 17) — consistent with the P0-4/Wave-4 persistence gates. - cell26 and cell23 remain on
python:3.9.18images (Wave 2, item 6). - The auto-merge bot path to
mainwithout required tests exists and has previously caused undeployed merges (July 6–8 backlog incident) — supporting Wave 1, items 1–2. deploy-all.ymljobs already requestid-token: write, so the WIF migration (CI-002) has no workflow-permission blocker.
Verdict: PLAN VALIDATED AND ADOPTED. Two precision annotations were folded into P0-4 and P0-8 above; neither changes any required fix, exit gate, wave, or merge-sequence item.