Migration — IntentIdentity re-keyed on (tenant_id, identity_id)

⛔ RETIRED — this migration has no target and must not be attempted

Owner decision, 2026-08-09: Neo4j is NOT being re-provisioned. The platform is intentionally staying on Firestore as its knowledge-graph backend, and the Aura host in the neo4j-uri secret is NXDOMAIN by choice. This document re-keys IntentIdentity in a Neo4j instance that will never be provisioned, so there is nothing to run it against: an operator following it will get DNS failures, not a migration. See docs/lii/RUNBOOK.md §1 "Cell 35 graph backend".

The readiness assertion in §5 is wrong under this decision. It tells an operator to expect backend: "neo4j" on GET /readyz; a correctly-configured Cell 35 now answers backend: "memory", backend_expected: "memory", degraded: false. The runbook carries the current contract — assert that, never this.

What is still true and still load-bearing: the composite key itself. Cell 35's in-memory backend implements the same (tenant_id, identity_id) MERGE semantics (graph_store.py NODE_KEY_PROPS), so the tenant-isolation property this migration exists to establish holds in the running deployment without any migration step. §1's analysis of why the single-key scheme was unsafe remains the record of that decision and is worth reading; §§2–6's operator procedure is not.

Kept rather than deleted so nobody re-derives the composite-key design from scratch, and so a future reversal of the owner decision starts from a reviewed plan.

Original status line (superseded by the banner above): written, not run. Not deployed, not live-verified. No Neo4j driver is available in this environment and this repository executes nothing against a database. Every statement below is statement-level only: it is the Cypher an operator should run, reviewed for shape and scope, not Cypher that has been executed or whose row counts have been observed. Treat the row counts in the examples as placeholders, not measurements.

Claim label: built, pre-benchmark.


1. What changed and why

IntentIdentity used to be keyed on identity_id alone, with tenant_id as an ordinary property inside the SET payload:

MERGE (n:IntentIdentity {identity_id: $key}) SET n += $props

Identity identifiers in this platform are deterministic by rule (blueprint §3 Cell 35 — hashed email, CRM id, any externally-minted string). Two tenants therefore routinely derive the same identifier for two different data subjects. Under the old key those two subjects shared one node, and tenant_id was last-writer-wins. Everything in GOVERNANCE.md §4.0 follows from that one line, including N-3: DETACH DELETE under tenant-a removed the node tenant-b also owned, taking tenant-b's relationships with it.

The key is now composite and tenant-first:

MERGE (n:IntentIdentity {tenant_id: $k0, identity_id: $k1}) SET n += $props

tenant_id is excluded from the SET payload, so a later write cannot re-stamp an existing node even if a caller supplies one. graph_store.py NODE_KEY_PROPS["IntentIdentity"] == ("tenant_id", "identity_id") is the single source of that fact; every MERGE, MATCH and DELETE derives its pattern from it.

The other three labels are unchanged and need no migration: IntentTopic is the governed taxonomy mirror and is deliberately shared by every tenant, while SignalAggregate (sha256(tenant|topic|signal_type)) and IntentExplanation (sha256(tenant|identity|topic)) already have the tenant inside their keys.


2. Detect affected nodes BEFORE running anything

2.1 Is there anything to split?

Under the old key an identifier that two tenants both wrote collapsed into one node, so the collision is not visible on the IntentIdentity nodes themselves — only one tenant_id survived. The evidence lives on the attached edges and on the explanations, which were never shared: every edge carries a mandatory tenant_id (REQUIRED_EDGE_METADATA) and every IntentExplanation carries the tenant that wrote it.

// (A) Identity nodes whose ATTACHED EDGES name more than one tenant.
// These are the shared nodes. Each must become one node per tenant.
MATCH (n:IntentIdentity)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS edge_tenants
WHERE size([t IN edge_tenants WHERE t IS NOT NULL]) > 1
RETURN n.identity_id       AS identity_id,
       n.tenant_id         AS stamped_tenant,
       edge_tenants        AS tenants
ORDER BY identity_id;
// (B) Identifiers whose EXPLANATIONS name more than one tenant. Same
// collision seen from the other side; run both, they can disagree when a
// tenant wrote scores but no explanation (or the reverse).
MATCH (e:IntentExplanation)
WITH e.identity_id AS identity_id, collect(DISTINCT e.tenant_id) AS tenants
WHERE size(tenants) > 1
RETURN identity_id, tenants
ORDER BY identity_id;
// (C) The single number to record before and after: how many identifiers are
// held under more than one tenant, by either signal.
MATCH (n:IntentIdentity)
OPTIONAL MATCH (n)-[r]-()
WITH n.identity_id AS identity_id,
     collect(DISTINCT r.tenant_id) + [n.tenant_id] AS ts
UNWIND ts AS t
WITH identity_id, collect(DISTINCT t) AS tenants
WHERE size([x IN tenants WHERE x IS NOT NULL]) > 1
RETURN count(*) AS identifiers_held_by_more_than_one_tenant;

2.2 Blocking preconditions

The forward migration must not proceed while either of these returns a non-zero count.

// (D) Untenanted identity nodes. The new erasure clause pins BOTH key
// properties and deliberately does NOT tolerate a NULL tenant — a
// NULL-tolerant identity clause is a match on identity_id alone again. If
// this returns > 0, decide an owner for each node (or delete it) FIRST;
// leaving them makes the strict clause under-delete.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
RETURN count(n) AS untenanted_identity_nodes;
// (E) Untenanted edges on identity nodes. These cannot be attributed to a
// tenant, so step 3 cannot decide which split node they belong to. Erasure
// still sweeps them (a blank edge tenant_id stays IN scope so a subject's own
// pre-stamping data is never under-deleted), but the SPLIT cannot place them.
MATCH (n:IntentIdentity)-[r]-()
WHERE r.tenant_id IS NULL OR r.tenant_id = ''
RETURN count(r) AS untenanted_edges;

Record the outputs of (C), (D) and (E). They are the before-numbers the verification in §5 is compared against.


3. Forward migration (operator-run)

Run against a backup or a restorable snapshot. Agents never run this; a human applies it (blueprint §9 human-gated rollout; Part 2 hard prohibition 6). Cell 35 should be scaled to zero, or its revision left on the pre-migration build, for the duration — see §4 for why the mixed states are unsafe.

Step 1 — drop the constraint that forbids the target state

DROP CONSTRAINT intent_identity_id_unique IF EXISTS;
DROP INDEX intent_identity_tenant IF EXISTS;

The old constraint requires identity_id to be globally unique, which is exactly what two tenants holding the same identifier must be allowed to violate. It has to go before the split, or step 3 fails on its first duplicate. The separate tenant_id index is dropped because the composite constraint's own index has tenant_id as its leading column and serves tenant-prefixed scans.

Step 2 — resolve the blocking preconditions

For each node from query (D), either assign the owning tenant or delete it. There is no automatic answer; the tenant is not recoverable from the node. If the attached edges agree on one tenant, that is the defensible choice:

// Assign an untenanted identity node the tenant its own edges all name.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
MATCH (n)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS ts
WHERE size([t IN ts WHERE t IS NOT NULL AND t <> '']) = 1
SET n.tenant_id = [t IN ts WHERE t IS NOT NULL AND t <> ''][0]
RETURN count(n) AS assigned;

Nodes that this does not resolve (no edges, or edges naming several tenants) must be decided by hand. Do not guess a tenant: a wrong stamp is a cross-tenant assignment written by the migration itself.

Step 3 — split each shared node, one node per tenant, re-pointing its edges

Run per identifier from query (A)/(B) — or as the loop below. It creates the missing per-tenant nodes and moves each edge to the node its own tenant_id names. apoc.refactor.to / from re-point a relationship without recreating it, preserving its properties and its identity.

// 3a. Create the missing per-tenant identity nodes.
MATCH (n:IntentIdentity)-[r]-()
WITH n, collect(DISTINCT r.tenant_id) AS ts
UNWIND [t IN ts WHERE t IS NOT NULL AND t <> ''] AS tenant
WITH n, tenant WHERE tenant <> n.tenant_id
MERGE (m:IntentIdentity {tenant_id: tenant, identity_id: n.identity_id})
  ON CREATE SET m.identity_kind = n.identity_kind,
                m.updated_at    = n.updated_at,
                m.migrated_from = 'single_key_split',
                m.migrated_at   = datetime()
RETURN count(m) AS nodes_created;
// 3b. Re-point OUTGOING edges to the node their own tenant_id names.
// Requires APOC. Without APOC, see 3b-alt.
MATCH (n:IntentIdentity)-[r]->(target)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CALL apoc.refactor.from(r, m) YIELD input, output
RETURN count(*) AS edges_repointed;
// 3c. Same for INCOMING edges (none are written today — every Cell 35 edge
// leaves the identity node — but the sweep must be symmetric or a hand-written
// edge is stranded).
MATCH (source)-[r]->(n:IntentIdentity)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CALL apoc.refactor.to(r, m) YIELD input, output
RETURN count(*) AS edges_repointed;
// 3b-alt. No APOC: recreate and delete, per relationship TYPE. Cell 35 writes
// exactly four types and the type cannot be parameterized, so this is four
// near-identical statements. HAS_INTENT shown; repeat verbatim for EXHIBITED,
// HAS_EXPLANATION and CITES (CITES leaves IntentExplanation, so it is not
// affected — three statements in practice).
MATCH (n:IntentIdentity)-[r:HAS_INTENT]->(target)
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
MATCH (m:IntentIdentity {tenant_id: r.tenant_id, identity_id: n.identity_id})
CREATE (m)-[r2:HAS_INTENT]->(target)
SET r2 = properties(r)
DELETE r
RETURN count(r2) AS edges_recreated;

SET r2 = properties(r) copies every property including the MERGE-pattern ones (horizon on HAS_INTENT), which is what keeps one edge per (identity, topic, horizon) after the move.

Step 4 — install the composite constraint

// Enterprise edition — ALSO enforces that both properties exist:
CREATE CONSTRAINT intent_identity_tenant_key IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE (n.tenant_id, n.identity_id) IS NODE KEY;

// Community / portable form:
CREATE CONSTRAINT intent_identity_tenant_key IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE (n.tenant_id, n.identity_id) IS UNIQUE;

This is the last statement of the forward migration on purpose. Its presence is the completion marker: _Neo4jBackend.verify_identity_key_schema refuses to serve unless intent_identity_tenant_key is present and intent_identity_id_unique is absent. Creating it first would tell the code the migration is finished while the split is still running.

Verify which form you got — the two are not equivalent, and the difference is whether a NULL tenant_id can reappear later:

SHOW CONSTRAINTS YIELD name, type, labelsOrTypes, properties
WHERE name = 'intent_identity_tenant_key'
RETURN name, type, labelsOrTypes, properties;

4. Compatibility — which combinations are safe

Code Data Verdict
New (composite key) Migrated Correct. The intended state.
New Un-migrated UNSAFE — and refused. See below.
Old (single key) Migrated UNSAFE — and NOT refused. Do not do this.
Old Un-migrated The pre-migration state. Carries N-3; correct only in the sense that it is what was there before.

New code against un-migrated data — unsafe, and it fails closed

A write for a second tenant MERGEs a new node, while that tenant's historical edges stay attached to the old shared node. Its own erasure then pins (tenant_id, identity_id), does not reach those stranded edges, and returns a clean receipt with counts that do not include them — silent under-deletion reported as a completed erasure, which is worse than an outage.

This combination is refused, not merely documented. _Neo4jBackend.__init__ calls verify_identity_key_schema(), which issues SHOW CONSTRAINTS YIELD name RETURN collect(name) AS names and raises GraphStoreError unless intent_identity_tenant_key is present and intent_identity_id_unique is absent. A deployed revision (K_SERVICE set) then refuses to boot per CONSTITUTION II.11 and traffic stays on the last healthy revision; a local process degrades loudly — /readyz 503, /health 200, and every write path refused 503. Tests: cell35/tests/test_tenant_scoped_identity.Neo4jSchemaGateTests.

The half-migrated state is caught by the same gate: leaving intent_identity_id_unique in place while adding the composite constraint forbids exactly the two-tenants-one-identifier state the new key needs, so the second tenant's MERGE raises a constraint violation. That surfaces as GraphStoreError → 503, which is fail-closed but is an outage; the gate catches it at boot instead.

Old code against migrated data — unsafe and NOT detected

This is the combination to plan around. A pre-migration Cell 35 has no schema gate, so nothing stops it. It issues MERGE (n:IntentIdentity {identity_id: $key}), which matches whichever of the now-several nodes the planner returns first and then SET n += $props re-stamps that node's tenant_id — reintroducing the shared node on top of migrated data, non-deterministically, one tenant at a time. Its erasure then DETACH DELETEs by identity_id with a mutable-property predicate, i.e. N-3, against data that has already been split.

The old code cannot be made to refuse — it has no gate to add without deploying new code, which defeats the point. The mitigation is procedural and must be part of the change plan:

  1. migrate after the new Cell 35 image is built and ready to deploy;
  2. do not leave a pre-migration revision able to take traffic — no traffic split, no rollback-to-previous-revision after the data is migrated (§6 covers the correct rollback: roll the DATA back first);
  3. Cell 33 is unaffected by the key change and does not need sequencing against it; its receipt tolerance for the superseded no_subject_data_foreign_identifier status stays for ordinary rolling deploys (cell33/ingest_cell/storage.py).

5. Verification — is the migration complete?

All four must hold. (V1)–(V3) are the completion proof; (V4) is the before/after comparison against §2.

// (V1) The composite constraint exists and the legacy one does not.
// This is exactly what the running code checks before it will serve.
SHOW CONSTRAINTS YIELD name
WITH collect(name) AS names
RETURN 'intent_identity_tenant_key' IN names  AS composite_present,
       NOT 'intent_identity_id_unique' IN names AS legacy_dropped;
// (V2) No identity node is untenanted. Must be 0 — the erasure clause pins
// both key properties and will not match a NULL tenant.
MATCH (n:IntentIdentity) WHERE n.tenant_id IS NULL OR n.tenant_id = ''
RETURN count(n) AS untenanted_identity_nodes;
// (V3) THE MIGRATION'S OWN POST-CONDITION: no identity node has an edge
// belonging to a tenant other than its own. Must be 0. A non-zero row here is
// a node the split missed, and it is precisely the shape that makes the new
// erasure under-delete.
MATCH (n:IntentIdentity)-[r]-()
WHERE r.tenant_id IS NOT NULL AND r.tenant_id <> '' AND r.tenant_id <> n.tenant_id
RETURN n.tenant_id AS node_tenant, r.tenant_id AS edge_tenant,
       n.identity_id AS identity_id, type(r) AS edge_type
LIMIT 50;
// (V4) Every identifier that query (C) reported as multi-tenant now has one
// node PER tenant. Compare the count against the (C) number recorded in §2.
MATCH (n:IntentIdentity)
WITH n.identity_id AS identity_id, count(*) AS nodes,
     collect(DISTINCT n.tenant_id) AS tenants
WHERE nodes > 1
RETURN count(*) AS identifiers_now_split,
       sum(CASE WHEN nodes = size(tenants) THEN 0 ELSE 1 END)
         AS identifiers_with_duplicate_tenant_nodes;

identifiers_with_duplicate_tenant_nodes must be 0 — a non-zero value means two nodes with the same (tenant_id, identity_id), which the composite constraint should have prevented and which indicates step 4 did not take.

Then, from a Cell 35 revision on the new code, an authenticated GET /readyz must answer 200 with degraded: false and backend: "neo4j". A 503 with degraded: true after a successful migration means the backend was constructed before the constraint existed, or the gate is still refusing — read the Cell 35 log line, it names which constraint it did not find.

⛔ RETIRED assertion — do not use. Per the 2026-08-09 owner decision there is no Neo4j, so backend: "neo4j" is unreachable and asserting it fails a correctly configured cell. The current contract is backend: "memory", backend_expected: "memory", degraded: false, HTTP 200 — see docs/lii/RUNBOOK.md §1, "The readiness assertion for Cell 35 — DECIDED". The paragraph above is kept only as the record of what the migration would have asserted.


6. Rollback

Roll the data back before the code, never the other way round. A rolled-back Cell 35 against migrated data is the undetected-unsafe combination in §4.

6.1 Preferred — restore the snapshot

Restore the pre-migration backup taken in §3. This is the only rollback that returns the graph to a state whose provenance is known: the merge below is lossy in a way that cannot be undone a second time.

6.2 If a snapshot is not available — merge the split nodes back

The split nodes are tagged migrated_from: 'single_key_split', so the ones this migration created can be identified and their edges returned:

// R1. Move every edge off a migration-created node back onto the surviving
// original (the one without the marker), then delete the created node.
MATCH (m:IntentIdentity {migrated_from: 'single_key_split'})
MATCH (n:IntentIdentity {identity_id: m.identity_id})
WHERE n.migrated_from IS NULL
MATCH (m)-[r]->(target)
CALL apoc.refactor.from(r, n) YIELD input
WITH DISTINCT m
MATCH (m)<-[r2]-(source)
CALL apoc.refactor.to(r2, m) YIELD input AS i2
RETURN count(DISTINCT m) AS nodes_to_delete;

// R2. Then remove the now-detached created nodes.
MATCH (m:IntentIdentity {migrated_from: 'single_key_split'})
WHERE NOT (m)--()
DELETE m;
// R3. Restore the old schema. The single-key constraint will FAIL if any
// identifier still has more than one node — resolve those first (R1/R2).
DROP CONSTRAINT intent_identity_tenant_key IF EXISTS;
CREATE CONSTRAINT intent_identity_id_unique IF NOT EXISTS
FOR (n:IntentIdentity) REQUIRE n.identity_id IS UNIQUE;
CREATE INDEX intent_identity_tenant IF NOT EXISTS
FOR (n:IntentIdentity) ON (n.tenant_id);

What rollback restores, stated plainly: it restores the shared node, and with it N-3 — a DETACH DELETE under one tenant again removing another tenant's relationships. Rolling back is a decision to re-accept a known tenant-isolation defect, and the erasure receipts issued while rolled back cannot be cited as evidence that one tenant's data was untouched.

Only after R1–R3 may a pre-migration Cell 35 revision be allowed to take traffic again.


7. Where the code says all this

Fact File
The key table (("tenant_id", "identity_id")) src/cells/cell35/graph_cell/graph_store.py — NODE_KEY_PROPS
Arity guard — a bare identity_id raises graph_store.py — node_key, identity_ref
Tenant excluded from the SET payload graph_store.py — _Neo4jBackend.merge_node
Erasure pins both key properties graph_store.py — _Neo4jBackend.delete_identity
The schema gate graph_store.py — verify_identity_key_schema
Constraint names graph_store.py — IDENTITY_KEY_CONSTRAINT, LEGACY_IDENTITY_CONSTRAINT
Applied schema src/cells/cell35/schema/intent_graph.cypher
N-3 regression tests src/cells/cell35/tests/test_tenant_scoped_identity.py
← All docsView source on GitHub →