ADR-LII-001 — Subject rights across the ORACLE / LII cells
- Status: Accepted (implemented; not live-verified)
- Date: 2026-08-09
- claim_label:
built, pre-benchmark - Deciders: Owner (decision a); implementing session (decisions b–d, within the owner's directive)
- Affects:
src/cells/cell33–cell36,src/shared/mizoki_intent/subject.py - Amended 2026-08-10 — the Neo4j premise is retired. This ADR was written on
2026-08-09 describing Cell 35 as a Neo4j-backed graph. Later the same day the owner
decided Neo4j is NOT being re-provisioned; the platform stays on Firestore as its
knowledge-graph backend, and Cell 35 serves from its in-memory backend
permanently (Firestore appears only as the optional durability journal noted
below — never as the serving backend). Every "Neo4j" in the text below
should be read as "the Cell 35 graph store", whose live implementation is now
in-process memory — ephemeral unless the optional Firestore durability journal
added by
5ac3455eis switched on, which requires an operator to setINTENT_GRAPH_DURABLE_PROJECTand is off by default. The decisions themselves — a per-cell erasure contract, the cascade, worst-outcome folding — are unaffected, because they were deliberately written against the store interface and not against Neo4j. Current operator truth:docs/lii/RUNBOOK.md§1 "Cell 35 graph backend". The body below is left as written, per ADR convention. - Supersedes / conflicts with: nothing. This ADR records decisions made on branch
claude/oracle-lii-build-bcljuc; the subject-rights code landed in commit82256da("feat(lii): GDPR subject access + erasure cascade for cells 33-36"). - Describes code later than
82256da.src/shared/mizoki_intent/subject.pywas amended on 2026-08-09 after this ADR was first written: thebefore == 0short-circuit was removed fromerase_with_receipt, andoverall_status([])was changed fromerasedtofailed. Decision (d) below describes the amended file (read 2026-08-09; 22483 bytes, md55fb48ca6…), not the state at82256da. Allsubject.pyline citations were recomputed against it.src/cells/cell34/tests/ test_subject_rights.pywas not unchanged — that same commit added 5 tests to it (20 → 25). The amendment is a pure addition: no pre-existing test in the file was removed or modified, so nothing was weakened to make the new behavior pass (25 passed, 2026-08-09). - Amended again 2026-08-09, after a second adversarial review. The
streaming-buffer branch now RECOUNTS before judging — BigQuery refuses DML at the
table level, so a subject with no rows was being told to retry in 90 minutes on a
409 whose receipt incoherently read
pendingwithremaining: 0; it is nowerased, because the guarantee is about the post-condition, not about whether a DELETE executed. The residual check isremaining != 0rather than> 0, so a store returning a negative count can no longer read as success. The docstring now states the guarantee's exact scope instead of implying the unconditional delete closes divergences the recount cannot see either.
Context
The ORACLE / LII platform was already built and on main before this session. It is
four Cloud Run cells under src/cells/cellNN/ (not services/cellNN_name/):
| Cell | Service | Role | SRPVDAL phase |
|---|---|---|---|
| 33 | intent-signal-ingest |
consent-gated signal + outcome ingestion into mizoki_intent |
SENSE |
| 34 | intent-scoring-api |
the Intent API: scores, cohorts, transitions, explain proxy, taxonomy | REASON / serving |
| 35 | intent-graph |
intent graph + explanation paths (written as Neo4j; in-memory since the 2026-08-09 retirement — see the amendment above) | REASON |
| 36 | intent-causal |
permanent holdouts, DR-Learner + CUPED estimators, incrementality | VALIDATE / DECIDE evidence |
(cell28 is an unrelated Sales Revenue Pipeline and has no part in this stack.)
Consent was already gated at ingest, satisfying the first half of CONSTITUTION.md
Art. II.6. The second half — "GDPR/CCPA subject access and erasure are honored
immediately" — had no implementation. Adding it touches all four cells, because each
owns different stores:
| Cell | Stores holding subject data |
|---|---|
| 33 | intent_signals, intent_outcomes |
| 34 | intent_scores, intent_transitions |
| 35 | the subject's subgraph in the Cell 35 graph store (Neo4j as written; in-memory in every current deployment) |
| 36 | intent_outcomes, intent_holdouts |
The forces in play:
- Art. II.6 is a non-negotiable, so the mechanism has to be fail-closed and auditable.
TRUTH.md5.4 andCONSTITUTION.mdIII.5 forbid smoothing a failure into a success — and BigQuery will refuse a DELETE against rows in the streaming buffer for up to ~90 minutes (AGENTS.md6.5), which means some erasures genuinely cannot complete when requested.- Every additional inter-cell caller means another service-account identity, another
ALLOWED_CALLER_SAentry, another IAM grant, another way to be wrong.
Decision (a) — Extend cells 33–36 in place; do not create new cells
Owner decision, this session. The subject-rights surface is added to the four existing cells rather than shipped as a new "privacy cell" (which would have been cell 37+).
Why.
- A privacy cell cannot honestly answer for stores it does not own. Erasure has to be verified by a recount against the actual store. A separate cell would have to be granted write access to every intent table and the Neo4j instance, then trust its own view of someone else's data. Erasure verification belongs where the data lives.
- Blast radius. A new cell means a new Cloud Run service, a new runtime SA, new IAM
grants on
mizoki_intentand Neo4j, a new deploy job, new smoke coverage — for a surface that is three routes. - The platform's stated posture (root
CLAUDE.md§3) prioritizes hardening and consolidation over new-parallel-path expansion, and.claude/rules/platform-services.mdsays "prefer hardening/consolidation over new parallel paths".
Consequences.
- Four services change instead of one, and all four must be redeployed for the feature
to be complete. The deploy workflow already deploys them individually
(
cellsinput), so this is a sequencing cost, not a structural one. - Each cell carries a little more code. Mitigated by decision (c): the shared logic is in one module and each cell supplies only its store adapters.
- Rejected alternative: a "privacy cell" or a BigQuery-only scheduled purge job. The purge job was rejected outright — it produces no per-request receipt, cannot see Neo4j, and gives the data subject nothing to be shown.
Decision (b) — Cell 34 is the cascade coordinator
POST /v1/intent/subject/{identity_id}:erase exists only on Cell 34
(src/cells/cell34/scoring_cell/main.py:293-310), which erases its own stores and then
fans out to Cells 33, 35 and 36
(src/cells/cell34/scoring_cell/orchestrator.py:763-800).
Why Cell 34 and not one of the others.
- It already holds the stack's only cross-cell client. Cell 34 calls Cell 35 through
INTENT_GRAPH_URLfor the explain proxy. Putting the fan-out anywhere else would mint a second inter-cell caller identity — a second SA needingroles/run.invokeron three services and membership in threeALLOWED_CALLER_SAallowlists. The code states this reasoning inline atorchestrator.py:567-571andmain.py:18-24. - Cell 34 is the primary serving API and the natural front door for an operator or a DSAR tool.
- Cell 34 is the one always-warm cell (
--min-instances=1), so the coordinating request does not pay three cold starts plus its own.
Why every cell still keeps a local DELETE. The coordinator is a convenience over a
uniform contract, not a privileged path. Each cell can be erased directly, which is what
makes the cascade auditable: the coordinator's receipts and a direct call to the peer
return the same receipt shape for the same store.
Fail-closed fan-out. An unset peer URL, an unreachable peer, and a peer that answers
without receipts are all failed receipts, never skipped stores
(orchestrator.py:693-711) — "this cell's copy of the subject may still exist, and the
cascade must say so (TRUTH.md 5.4)". A peer receipt carrying an unrecognized status is
coerced to failed (orchestrator.py:664-687): a peer cannot talk its way into a
completed erasure.
Consequences.
- Cell 34 needs
INTENT_INGEST_URL,INTENT_GRAPH_URL,INTENT_CAUSAL_URL. These are operator-set and currently OPEN (docs/lii/RUNBOOK.md§0 item 2) — until they are set, a cascade returns 502 withfailedreceipts for cells 33 and 36, i.e. a DSAR would not fully erase.GET /healthexposescascade_peers_configuredso the condition is visible before a request arrives (orchestrator.py:818-825). - Cell 34 is a single point of coordination. Accepted: a direct
DELETEper cell is always available as the manual fallback. - Rejected alternative: peer-to-peer fan-out from whichever cell receives the request. It multiplies caller identities by four and makes "who erased what" ambiguous in the audit.
Decision (c) — Contract shape: per-cell local endpoints + one coordinator
GET /v1/intent/subject/{identity_id} # all four cells — this cell's stores
DELETE /v1/intent/subject/{identity_id} # all four cells — this cell's stores
POST /v1/intent/subject/{identity_id}:erase # Cell 34 only — the cascade
Why access is per-cell and deliberately NOT a fan-out. Each store answers for itself,
so no cell can claim to speak for another's contents
(orchestrator.py:581-586). A full subject file is assembled by calling all four GETs.
An aggregating read would have to describe data it cannot see failing, which is exactly
the honesty failure the erasure design exists to prevent.
Why one shared module. All receipt logic lives in
src/shared/mizoki_intent/subject.py; each cell supplies only count / delete /
select callables for its own stores. That is what makes the receipts byte-identical
in shape across four cells with three different backends (BigQuery, Neo4j, in-memory),
which is the property the coordinator relies on.
Why the subject-rights routes do NOT use the §6.1 response envelope. Every other
Cell 34 data route wraps its payload in {data, model_version, calibration_version,
claim_label, consent_basis, tenant_id, generated_at}. The receipt shape is a fixed
cross-cell contract that Cells 33/35/36 return identically; wrapping only Cell 34's copy
would break that symmetry (main.py:26-28).
Status semantics as part of the contract (subject.py:53-71):
| Per-store status | HTTP | Meaning |
|---|---|---|
erased |
200 | verified gone — the recount returned 0 |
pending_streaming_buffer |
409 | rows still present; retry after ~90 min. Not a success |
failed |
502 | store unreachable / uncountable / undeletable, or peer unconfigured |
The overall status is the worst across stores. Only a fully verified erasure is 2xx — a partial cascade is deliberately non-2xx so no caller can mistake it for a completed one.
Idempotency over 404. Erasing an unknown or already-erased identity is a zero-count
success, never a 404 — a 404 would leak whether an identity exists to any caller who
can guess an id (subject.py:29-32). The zero-count success is not short-circuited
from the pre-count: the delete is issued and the recount runs anyway, so even the
idempotent case reports a measured result. See decision (d), "The recount guarantee,
stated exactly".
Consequences.
- A DSAR tool must call four GETs to assemble a full subject file. Accepted; the alternative is a cell asserting things about stores it cannot verify.
- Adding a fifth cell later means adding one entry to
CASCADE_PEERS(orchestrator.py:572-576) and one env var — no protocol change.
Decision (d) — Erasure uses count → delete → RE-COUNT
erase_with_receipt (src/shared/mizoki_intent/subject.py:133-201) counts the
subject's rows, deletes them, then counts again, and reports what actually happened.
The delete and the recount are unconditional — see "The recount guarantee, stated
exactly" below for what that does and does not buy.
Why the recount is not optional.
- BigQuery's streaming buffer makes "the DELETE returned without error" a lie. Rows
streamed within roughly the last 90 minutes are not touchable by DML (
AGENTS.md6.5). A DELETE issued seconds after ingest fails — sometimes by raising, and the honest outcome in either case is "rows are still there". - Neo4j and in-memory backends have their own ways to under-delete (a partially
matched
DETACH DELETE, a concurrent write). The recount is backend-agnostic: it asks the store what it now holds, rather than trusting the delete's return value. TRUTH.md5.4 /CONSTITUTION.mdIII.5: failures are reported with their evidence, never smoothed into success. A receipt that sayserasedwithout a verifying observation would be exactly that smoothing.
Every branch is honest, including the ones that fail (subject.py:157-201):
| Situation | Receipt | Rationale |
|---|---|---|
| Pre-count fails | failed |
cannot even see the store |
| Pre-count is 0 | the delete is still issued and the recount still runs; a recount of 0 gives erased with zero counts |
idempotent zero-count success — but earned by a post-delete observation, not assumed from the pre-count |
| Delete raises a streaming-buffer error | pending_streaming_buffer + retry_after_seconds: 5400 |
retryable incompleteness, not a failure and never a success |
| Delete raises anything else | failed |
|
| Post-count fails | failed, detail "delete issued but could not be verified" |
an erasure that cannot be verified is never claimed |
| Post-count > 0 | pending_streaming_buffer with the surviving count |
rows survived; say so |
| Post-count == 0 | erased with the erased count |
the only success |
The recount guarantee, stated exactly
An earlier revision of this ADR asserted "there is no code path that reports erased
without a verifying recount" two lines below a table row that described exactly such a
path — Pre-count is 0 → erased, which returned before issuing the delete and before
recounting. That was a self-contradiction, and the guarantee as written was false.
It is now true, because the code changed. Verified by reading
src/shared/mizoki_intent/subject.py on 2026-08-09 (file md5 5fb48ca6…, 22483 bytes):
- The delete is unconditional. The
if before == 0: return …short-circuit is gone. Control flows from the pre-count straight intodelete_fn()(subject.py:157-164), so a zero pre-count no longer skips the DML. The code states the reason inline (subject.py:148-155): "the count predicate and the store's real contents can diverge (a row whose tenant_id was coerced, a NULL key, a legacy row the WHERE clause misses), and short-circuiting onbefore == 0would then reporterasedwhile the subject's row is still sitting in the table." - The recount is unconditional. Every non-raising delete falls through to a second
count_fn()(subject.py:182-192), commented "ALWAYS recount, including after a zero-row pre-count: the recount is what makes 'erased' a measurement rather than an assumption." - Therefore: every
erasedreceipt in this module — including the zero-count one — is now backed by a post-delete observation that returned 0 (subject.py:194-201). There is no argument, flag, or caller-supplied callable that skips it. - Idempotency is unchanged. Erasing an unknown or already-erased identity is still a zero-count success, never a 404. The cost is one extra idempotent DML per DSAR, which decision (d) already accepts as cheap.
A related fail-open hole was closed with it. overall_status([]) previously returned
erased — the fold's identity element — so an empty receipt list would have rendered
as 200 erased: an erasure asserted across zero stores. overall_status now returns
failed on an empty list (subject.py:204-221), and erasure_response attaches the
detail "no store receipts — nothing was verified, so this erasure is unproven"
(subject.py:590-597). An empty receipt list is now a non-2xx failure. Note this was
never reachable from a route — all four cells build fixed, non-empty receipt lists and
_erase_peer returns a failed receipt rather than an empty list on every non-answer
(src/cells/cell34/scoring_cell/orchestrator.py:701-760) — so closing it removed a latent
contract defect, not a live one.
The module docstring's sentence at subject.py:25-27 now matches the code.
Why DELETE and not UPDATE/tombstone. Bitemporal event columns are append-only facts
and are never rewritten in place (AGENTS.md 6.2), so erasure removes the rows
themselves (subject.py:515-522). A tombstone column would leave the subject's data in
the table.
Why the audit stores a salted hash. Erasure is a governed action and must be
auditable, but "an erasure audit that stored the erased identifier would defeat the
erasure it records" (subject.py:34-36). The audit row carries
subject_ref = sha256(salt | tenant_id | identity_id) (subject.py:312-323) — stable,
so repeat DSARs for one subject stay correlatable, and one-way, so the audit cannot
re-identify. The row schema is a hard allowlist re-applied in code
(subject.py:342-345, 396), and the DDL carries no topic_id, no payload, no signal
content (src/cells/cell33/schema/intent_bigquery_ddl.sql:184-216).
Why audit failures do not block erasure. The receipt returned to the caller is the
primary record; the audit table is the durable governance copy. An audit insert error is
logged and swallowed (subject.py:397-408) so a missing table cannot stop a subject
exercising their rights. The consequence is the reverse dependency recorded as
docs/lii/RUNBOOK.md §0 item 1: until the table-9 DDL is applied, audit inserts log an
error and fall through.
Consequences.
- Every erasure costs at least two extra store reads. Accepted; DSARs are rare and correctness dominates.
- Callers must handle 409 as a real outcome — an erasure that is genuinely not finished yet. This is documented in the runbook's DSAR procedure.
- Rejected alternative: report
erasedon a non-raising DELETE. Simpler, faster, and wrong under the streaming buffer — it would produce a receipt claiming a completed erasure while the subject's rows were still queryable.
Compliance and verification status
- Implemented, NOT live-verified. The streaming-buffer path is fault-injected in
tests (
src/cells/cell34/tests/test_subject_rights.py:83, 329, 469) and has never been observed against real BigQuery. No claim in this ADR is a live-verified claim. - Neither live backend path is executed by any test — SQL and Cypher alike. Neither
neo4jnorgoogle-cloud-bigqueryis installed in the environment these tests run in (bothimportcalls raiseModuleNotFoundError; verified 2026-08-09), and every live path is behind a lazy, function-local import taken only when a real client is present. Nothing insrc/cells/cell33–cell36/tests/referencesbq_select_subject_rows,bq_count_subject_rows,bq_delete_subject_rows,_count_subject_live,_causal_join_liveor_Neo4jBackend(grep, zero hits). So the BigQueryDELETE(subject.py:511-536) and the Neo4jDETACH DELETE(src/cells/cell35/graph_cell/graph_store.py:500-507) have never been executed, never been parsed at runtime, and never been type-checked against a live driver. The whole of decision (d) is verified against in-memory doubles only. Status: implemented, unexecuted on the live path. - Amended 2026-08-10. Both halves of that bullet have moved. The BigQuery
path is now live-verified: a real DSAR round-trip ran it on 2026-08-10, including
the streaming-buffer 409 and a clean 200 on retry (evidence:
.claude/memory/inbox/2026-08.md; 39 rows inintent_erasure_audit). It remains true that no test executes it. The Neo4j path is not merely unexecuted but unreachable — Neo4j was retired on 2026-08-09, no deployment setsNEO4J_URI, so_Neo4jBackendis never constructed and theDETACH DELETEis dead code. - The "vector index" limb of the erasure cascade is NOT-APPLICABLE, not satisfied.
Outside descriptions of this stack say erasure "cascades BQ + Neo4j + vector index"
(e.g.
skills/miz-oki-platform-expert/SKILL.md:1449). There is no vector index, vector store, or embedding of any kind in Cells 33–36 orsrc/shared/mizoki_intent/— a case-insensitive grep forvector index|vector store|vector search|vector db| embedding|faiss|pinecone|weaviate|qdrant|milvus|chroma|pgvectoracross those paths returns zero hits (verified 2026-08-09). The cascade covers the two stores that exist: BigQuery and the Cell 35 graph store (written as Neo4j; in-memory since the 2026-08-09 retirement). It does not cover a vector index because there is nothing to cover. If one is ever added to this stack, the cascade does not extend to it automatically and MUST be extended — a new store means a newerase_with_receiptcall, a new receipt line, and a new count/delete pair, or erasure silently stops being complete. - Test coverage added with the change:
test_subject_rights.pyin all four cells (cell33, cell34, cell35, cell36) — covering per-cell access/erasure, idempotency, tenant scoping, auth, streaming-buffer pending, partial-failure non-2xx, unconfigured and unreachable peers, unknown-peer-status coercion, deny-list scrubbing on read, and honest health reporting of peer configuration. - No governance gate was weakened. Erasing a Cell 36 holdout assignment does not
open a probabilistic path: assignment is a deterministic hash of the unit id and the
deterministic-only gate still refuses probabilistic identities
(
src/cells/cell36/causal_cell/orchestrator.py:406-410, asserted bytest_probabilistic_identity_still_refused_after_erasure). - Open operator dependencies are recorded in
docs/lii/RUNBOOK.md§0 and are not code changes: apply theintent_erasure_auditDDL; set the two cascade peer URLs on Cell 34; setINTENT_SUBJECT_SALTper deployment. - Related:
docs/lii/GOVERNANCE.md(the enforced gates, with citations),docs/INTENT_API_PLATFORM_BLUEPRINT.md§8 / §9 / Appendix A,docs/reports/INTENT_PHASE_A_GATE_REPORT_2026-08-01.md(the ranking gate — failed and open).