Migration — IntentIdentity re-keyed on (tenant_id, identity_id)
⛔ RETIRED — this migration has no target and must not be attempted
Owner decision, 2026-08-09: Neo4j is NOT being re-provisioned. The platform is intentionally staying on Firestore as its knowledge-graph backend, and the Aura host in the
neo4j-urisecret is NXDOMAIN by choice. This document re-keysIntentIdentityin a Neo4j instance that will never be provisioned, so there is nothing to run it against: an operator following it will get DNS failures, not a migration. Seedocs/lii/RUNBOOK.md§1 "Cell 35 graph backend".The readiness assertion in §5 is wrong under this decision. It tells an operator to expect
backend: "neo4j"onGET /readyz; a correctly-configured Cell 35 now answersbackend: "memory",backend_expected: "memory",degraded: false. The runbook carries the current contract — assert that, never this.What is still true and still load-bearing: the composite key itself. Cell 35's in-memory backend implements the same
(tenant_id, identity_id)MERGE semantics (graph_store.pyNODE_KEY_PROPS), so the tenant-isolation property this migration exists to establish holds in the running deployment without any migration step. §1's analysis of why the single-key scheme was unsafe remains the record of that decision and is worth reading; §§2–6's operator procedure is not.Kept rather than deleted so nobody re-derives the composite-key design from scratch, and so a future reversal of the owner decision starts from a reviewed plan.
Original status line (superseded by the banner above): written, not run. Not deployed, not live-verified. No Neo4j driver is available in this environment and this repository executes nothing against a database. Every statement below is statement-level only: it is the Cypher an operator should run, reviewed for shape and scope, not Cypher that has been executed or whose row counts have been observed. Treat the row counts in the examples as placeholders, not measurements.
Claim label: built, pre-benchmark.
1. What changed and why
IntentIdentity used to be keyed on identity_id alone, with tenant_id as
an ordinary property inside the SET payload:
MERGE (n:IntentIdentity {identity_id: $key}) SET n += $props
Identity identifiers in this platform are deterministic by rule (blueprint
§3 Cell 35 — hashed email, CRM id, any externally-minted string). Two tenants
therefore routinely derive the same identifier for two different data
subjects. Under the old key those two subjects shared one node, and
tenant_id was last-writer-wins. Everything in GOVERNANCE.md §4.0 follows
from that one line, including N-3: DETACH DELETE under tenant-a removed the
node tenant-b also owned, taking tenant-b's relationships with it.
The key is now composite and tenant-first:
MERGE (n:IntentIdentity {tenant_id: $k0, identity_id: $k1}) SET n += $props
tenant_id is excluded from the SET payload, so a later write cannot
re-stamp an existing node even if a caller supplies one. graph_store.py
NODE_KEY_PROPS["IntentIdentity"] == ("tenant_id", "identity_id") is the
single source of that fact; every MERGE, MATCH and DELETE derives its pattern
from it.
The other three labels are unchanged and need no migration: IntentTopic is
the governed taxonomy mirror and is deliberately shared by every tenant,
while SignalAggregate (sha256(tenant|topic|signal_type)) and
IntentExplanation (sha256(tenant|identity|topic)) already have the tenant
inside their keys.
2. Detect affected nodes BEFORE running anything
2.1 Is there anything to split?
Under the old key an identifier that two tenants both wrote collapsed into one
node, so the collision is not visible on the IntentIdentity nodes
themselves — only one tenant_id survived. The evidence lives on the
attached edges and on the explanations, which were never shared: every edge
carries a mandatory tenant_id (REQUIRED_EDGE_METADATA) and every
IntentExplanation carries the tenant that wrote it.
// (A) Identity nodes whose ATTACHED EDGES name more than one tenant.
// These are the shared nodes. Each must become one node per tenant.
MATCH (n:IntentIdentity)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS edge_tenants
WHERE size([t IN edge_tenants WHERE t IS NOT NULL]) > 1
RETURN n.identity_id AS identity_id,
n.tenant_id AS stamped_tenant,
edge_tenants AS tenants
ORDER BY identity_id;
// (B) Identifiers whose EXPLANATIONS name more than one tenant. Same
// collision seen from the other side; run both, they can disagree when a
// tenant wrote scores but no explanation (or the reverse).
MATCH (e:IntentExplanation)
WITH e.identity_id AS identity_id, collect(DISTINCT e.tenant_id) AS tenants
WHERE size(tenants) > 1
RETURN identity_id, tenants
ORDER BY identity_id;
// (C) The single number to record before and after: how many identifiers are
// held under more than one tenant, by either signal.
MATCH (n:IntentIdentity)
OPTIONAL MATCH (n)-[r]-()
WITH n.identity_id AS identity_id,
collect(DISTINCT r.tenant_id) + [n.tenant_id] AS ts
UNWIND ts AS t
WITH identity_id, collect(DISTINCT t) AS tenants
WHERE size([x IN tenants WHERE x IS NOT NULL]) > 1
RETURN count(*) AS identifiers_held_by_more_than_one_tenant;
2.2 Blocking preconditions
The forward migration must not proceed while either of these returns a non-zero count.
// (D) Untenanted identity nodes. The new erasure clause pins BOTH key
// properties and deliberately does NOT tolerate a NULL tenant — a
// NULL-tolerant identity clause is a match on identity_id alone again. If
// this returns > 0, decide an owner for each node (or delete it) FIRST;
// leaving them makes the strict clause under-delete.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
RETURN count(n) AS untenanted_identity_nodes;
// (E) Untenanted edges on identity nodes. These cannot be attributed to a
// tenant, so step 3 cannot decide which split node they belong to. Erasure
// still sweeps them (a blank edge tenant_id stays IN scope so a subject's own
// pre-stamping data is never under-deleted), but the SPLIT cannot place them.
MATCH (n:IntentIdentity)-[r]-()
WHERE r.tenant_id IS NULL OR r.tenant_id = ''
RETURN count(r) AS untenanted_edges;
Record the outputs of (C), (D) and (E). They are the before-numbers the verification in §5 is compared against.
3. Forward migration (operator-run)
Run against a backup or a restorable snapshot. Agents never run this; a human applies it (blueprint §9 human-gated rollout; Part 2 hard prohibition 6). Cell 35 should be scaled to zero, or its revision left on the pre-migration build, for the duration — see §4 for why the mixed states are unsafe.
Step 1 — drop the constraint that forbids the target state
DROP CONSTRAINT intent_identity_id_unique IF EXISTS;
DROP INDEX intent_identity_tenant IF EXISTS;
The old constraint requires identity_id to be globally unique, which is
exactly what two tenants holding the same identifier must be allowed to
violate. It has to go before the split, or step 3 fails on its first
duplicate. The separate tenant_id index is dropped because the composite
constraint's own index has tenant_id as its leading column and serves
tenant-prefixed scans.
Step 2 — resolve the blocking preconditions
For each node from query (D), either assign the owning tenant or delete it. There is no automatic answer; the tenant is not recoverable from the node. If the attached edges agree on one tenant, that is the defensible choice:
// Assign an untenanted identity node the tenant its own edges all name.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
MATCH (n)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS ts
WHERE size([t IN ts WHERE t IS NOT NULL AND t <> '']) = 1
SET n.tenant_id = [t IN ts WHERE t IS NOT NULL AND t <> ''][0]
RETURN count(n) AS assigned;
Nodes that this does not resolve (no edges, or edges naming several tenants) must be decided by hand. Do not guess a tenant: a wrong stamp is a cross-tenant assignment written by the migration itself.
Step 3 — split each shared node, one node per tenant, re-pointing its edges
Run per identifier from query (A)/(B) — or as the loop below. It creates the
missing per-tenant nodes and moves each edge to the node its own tenant_id
names. apoc.refactor.to / from re-point a relationship without recreating
it, preserving its properties and its identity.
// 3a. Create the missing per-tenant identity nodes.
MATCH (n:IntentIdentity)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS ts
UNWIND [t IN ts WHERE t IS NOT NULL AND t <> ''] AS tenant
WITH n, tenant WHERE tenant <> n.tenant_id
MERGE (m:IntentIdentity {tenant_id: tenant, identity_id: n.identity_id})
ON CREATE SET m.identity_kind = n.identity_kind,
m.updated_at = n.updated_at,
m.migrated_from = 'single_key_split',
m.migrated_at = datetime()
RETURN count(m) AS nodes_created;
// 3b. Re-point OUTGOING edges to the node their own tenant_id names.
// Requires APOC. Without APOC, see 3b-alt.
MATCH (n:IntentIdentity)-[r]->(target)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CALL apoc.refactor.from(r, m) YIELD input, output
RETURN count(*) AS edges_repointed;
// 3c. Same for INCOMING edges (none are written today — every Cell 35 edge
// leaves the identity node — but the sweep must be symmetric or a hand-written
// edge is stranded).
MATCH (source)-[r]->(n:IntentIdentity)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CALL apoc.refactor.to(r, m) YIELD input, output
RETURN count(*) AS edges_repointed;
// 3b-alt. No APOC: recreate and delete, per relationship TYPE. Cell 35 writes
// exactly four types and the type cannot be parameterized, so this is four
// near-identical statements. HAS_INTENT shown; repeat verbatim for EXHIBITED,
// HAS_EXPLANATION and CITES (CITES leaves IntentExplanation, so it is not
// affected — three statements in practice).
MATCH (n:IntentIdentity)-[r:HAS_INTENT]->(target)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CREATE (m)-[r2:HAS_INTENT]->(target)
SET r2 = properties(r)
DELETE r
RETURN count(r2) AS edges_recreated;
SET r2 = properties(r)copies every property including the MERGE-pattern ones (horizononHAS_INTENT), which is what keeps one edge per(identity, topic, horizon)after the move.
Step 4 — install the composite constraint
// Enterprise edition — ALSO enforces that both properties exist:
CREATE CONSTRAINT intent_identity_tenant_key IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE (n.tenant_id, n.identity_id) IS NODE KEY;
// Community / portable form:
CREATE CONSTRAINT intent_identity_tenant_key IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE (n.tenant_id, n.identity_id) IS UNIQUE;
This is the last statement of the forward migration on purpose. Its
presence is the completion marker: _Neo4jBackend.verify_identity_key_schema
refuses to serve unless intent_identity_tenant_key is present and
intent_identity_id_unique is absent. Creating it first would tell the code
the migration is finished while the split is still running.
Verify which form you got — the two are not equivalent, and the difference is
whether a NULL tenant_id can reappear later:
SHOW CONSTRAINTS YIELD name, type, labelsOrTypes, properties
WHERE name = 'intent_identity_tenant_key'
RETURN name, type, labelsOrTypes, properties;
4. Compatibility — which combinations are safe
| Code | Data | Verdict |
|---|---|---|
| New (composite key) | Migrated | Correct. The intended state. |
| New | Un-migrated | UNSAFE — and refused. See below. |
| Old (single key) | Migrated | UNSAFE — and NOT refused. Do not do this. |
| Old | Un-migrated | The pre-migration state. Carries N-3; correct only in the sense that it is what was there before. |
New code against un-migrated data — unsafe, and it fails closed
A write for a second tenant MERGEs a new node, while that tenant's
historical edges stay attached to the old shared node. Its own erasure
then pins (tenant_id, identity_id), does not reach those stranded edges, and
returns a clean receipt with counts that do not include them — silent
under-deletion reported as a completed erasure, which is worse than an
outage.
This combination is refused, not merely documented.
_Neo4jBackend.__init__ calls verify_identity_key_schema(), which issues
SHOW CONSTRAINTS YIELD name RETURN collect(name) AS names and raises
GraphStoreError unless intent_identity_tenant_key is present and
intent_identity_id_unique is absent. A deployed revision
(K_SERVICE set) then refuses to boot per CONSTITUTION II.11 and traffic
stays on the last healthy revision; a local process degrades loudly —
/readyz 503, /health 200, and every write path refused 503. Tests:
cell35/tests/test_tenant_scoped_identity.Neo4jSchemaGateTests.
The half-migrated state is caught by the same gate: leaving
intent_identity_id_unique in place while adding the composite constraint
forbids exactly the two-tenants-one-identifier state the new key needs, so
the second tenant's MERGE raises a constraint violation. That surfaces as
GraphStoreError → 503, which is fail-closed but is an outage; the gate
catches it at boot instead.
Old code against migrated data — unsafe and NOT detected
This is the combination to plan around. A pre-migration Cell 35 has no
schema gate, so nothing stops it. It issues
MERGE (n:IntentIdentity {identity_id: $key}), which matches whichever of
the now-several nodes the planner returns first and then SET n += $props
re-stamps that node's tenant_id — reintroducing the shared node on top of
migrated data, non-deterministically, one tenant at a time. Its erasure then
DETACH DELETEs by identity_id with a mutable-property predicate, i.e.
N-3, against data that has already been split.
The old code cannot be made to refuse — it has no gate to add without deploying new code, which defeats the point. The mitigation is procedural and must be part of the change plan:
- migrate after the new Cell 35 image is built and ready to deploy;
- do not leave a pre-migration revision able to take traffic — no traffic split, no rollback-to-previous-revision after the data is migrated (§6 covers the correct rollback: roll the DATA back first);
- Cell 33 is unaffected by the key change and does not need sequencing
against it; its receipt tolerance for the superseded
no_subject_data_foreign_identifierstatus stays for ordinary rolling deploys (cell33/ingest_cell/storage.py).
5. Verification — is the migration complete?
All four must hold. (V1)–(V3) are the completion proof; (V4) is the
before/after comparison against §2.
// (V1) The composite constraint exists and the legacy one does not.
// This is exactly what the running code checks before it will serve.
SHOW CONSTRAINTS YIELD name
WITH collect(name) AS names
RETURN 'intent_identity_tenant_key' IN names AS composite_present,
NOT 'intent_identity_id_unique' IN names AS legacy_dropped;
// (V2) No identity node is untenanted. Must be 0 — the erasure clause pins
// both key properties and will not match a NULL tenant.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
RETURN count(n) AS untenanted_identity_nodes;
// (V3) THE MIGRATION'S OWN POST-CONDITION: no identity node has an edge
// belonging to a tenant other than its own. Must be 0. A non-zero row here is
// a node the split missed, and it is precisely the shape that makes the new
// erasure under-delete.
MATCH (n:IntentIdentity)-[r]-()
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
RETURN n.tenant_id AS node_tenant, r.tenant_id AS edge_tenant,
n.identity_id AS identity_id, type(r) AS edge_type
LIMIT 50;
// (V4) Every identifier that query (C) reported as multi-tenant now has one
// node PER tenant. Compare the count against the (C) number recorded in §2.
MATCH (n:IntentIdentity)
WITH n.identity_id AS identity_id, count(*) AS nodes,
collect(DISTINCT n.tenant_id) AS tenants
WHERE nodes > 1
RETURN count(*) AS identifiers_now_split,
sum(CASE WHEN nodes = size(tenants) THEN 0 ELSE 1 END)
AS identifiers_with_duplicate_tenant_nodes;
identifiers_with_duplicate_tenant_nodes must be 0 — a non-zero value means
two nodes with the same (tenant_id, identity_id), which the composite
constraint should have prevented and which indicates step 4 did not take.
Then, from a Cell 35 revision on the new code, an authenticated
GET /readyz must answer 200 with degraded: false and backend: "neo4j".
A 503 with degraded: true after a successful migration means the backend
was constructed before the constraint existed, or the gate is still refusing —
read the Cell 35 log line, it names which constraint it did not find.
⛔ RETIRED assertion — do not use. Per the 2026-08-09 owner decision there is no Neo4j, so
backend: "neo4j"is unreachable and asserting it fails a correctly configured cell. The current contract isbackend: "memory",backend_expected: "memory",degraded: false, HTTP 200 — seedocs/lii/RUNBOOK.md§1, "The readiness assertion for Cell 35 — DECIDED". The paragraph above is kept only as the record of what the migration would have asserted.
6. Rollback
Roll the data back before the code, never the other way round. A rolled-back Cell 35 against migrated data is the undetected-unsafe combination in §4.
6.1 Preferred — restore the snapshot
Restore the pre-migration backup taken in §3. This is the only rollback that returns the graph to a state whose provenance is known: the merge below is lossy in a way that cannot be undone a second time.
6.2 If a snapshot is not available — merge the split nodes back
The split nodes are tagged migrated_from: 'single_key_split', so the ones
this migration created can be identified and their edges returned:
// R1. Move every edge off a migration-created node back onto the surviving
// original (the one without the marker), then delete the created node.
MATCH (m:IntentIdentity {migrated_from: 'single_key_split'})
MATCH (n:IntentIdentity {identity_id: m.identity_id})
WHERE n.migrated_from IS NULL
MATCH (m)-[r]->(target)
CALL apoc.refactor.from(r, n) YIELD input
WITH DISTINCT m
MATCH (m)<-[r2]-(source)
CALL apoc.refactor.to(r2, m) YIELD input AS i2
RETURN count(DISTINCT m) AS nodes_to_delete;
// R2. Then remove the now-detached created nodes.
MATCH (m:IntentIdentity {migrated_from: 'single_key_split'})
WHERE NOT (m)--()
DELETE m;
// R3. Restore the old schema. The single-key constraint will FAIL if any
// identifier still has more than one node — resolve those first (R1/R2).
DROP CONSTRAINT intent_identity_tenant_key IF EXISTS;
CREATE CONSTRAINT intent_identity_id_unique IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE n.identity_id IS UNIQUE;
CREATE INDEX intent_identity_tenant IF NOT EXISTS
FOR (n:IntentIdentity) ON (n.tenant_id);
What rollback restores, stated plainly: it restores the shared node, and
with it N-3 — a DETACH DELETE under one tenant again removing another
tenant's relationships. Rolling back is a decision to re-accept a known
tenant-isolation defect, and the erasure receipts issued while rolled back
cannot be cited as evidence that one tenant's data was untouched.
Only after R1–R3 may a pre-migration Cell 35 revision be allowed to take traffic again.
7. Where the code says all this
| Fact | File |
|---|---|
The key table (("tenant_id", "identity_id")) |
src/cells/cell35/graph_cell/graph_store.py — NODE_KEY_PROPS |
Arity guard — a bare identity_id raises |
graph_store.py — node_key, identity_ref |
Tenant excluded from the SET payload |
graph_store.py — _Neo4jBackend.merge_node |
| Erasure pins both key properties | graph_store.py — _Neo4jBackend.delete_identity |
| The schema gate | graph_store.py — verify_identity_key_schema |
| Constraint names | graph_store.py — IDENTITY_KEY_CONSTRAINT, LEGACY_IDENTITY_CONSTRAINT |
| Applied schema | src/cells/cell35/schema/intent_graph.cypher |
| N-3 regression tests | src/cells/cell35/tests/test_tenant_scoped_identity.py |