Marketing Journey Knowledge Graph Architecture (2026)

Purpose

This document integrates the latest marketing data pipeline patterns into the MIZOKI stack so raw ad/email/web signals are transformed into a governed, ontology-driven customer journey knowledge graph.

Target Reference Architecture

Marketing APIs / Webhooks
   ↓
Connector Layer
   ↓
Event Stream
   ↓
Schema Normalization
   ↓
Identity Resolution Engine
   ↓
Gemini Semantic Structuring
   ↓
Ontology Mapping
   ↓
Knowledge Graph
   ↓
AI Decision Agents

1) Source Extraction Patterns

Meta / Facebook Marketing API

Recommended ingestion mode: async + streaming hybrid.

  1. Trigger asynchronous Insights report jobs.
  2. Poll report status and fetch paginated result sets.
  3. Stream normalized records to the event bus.
POST /act_{account_id}/insights
fields=impressions,clicks,spend,actions
time_increment=1
level=ad

Implementation guidelines

Query standard: GAQL with fact/dimension separation.

Fact tables:

Dimension tables:

Implementation guidelines

Email engagement APIs

Open events should be treated as weak signals. Prioritize:

Reference flow:

ESP webhook → event queue → engagement classifier

Programmatic / DSP logs

Prefer log-level exports and retain raw event streams before rollups.

Common event types:

Web / app tracking

Prefer server-side tracking with stable identifiers:

2) Event Normalization Contract

All inbound events should be transformed into the canonical schema before enrichment and graph writes.

event_id
event_type
event_timestamp
actor_id
session_id
source_system
channel
campaign_id
creative_id
properties
value

Example:

event_type: ad_click
channel: google_ads
campaign_id: 812
actor_id: anon_73
landing_url: /product/shoes

3) Gemini Semantic Structuring Layer

Apply LLM-assisted enrichment after schema validation:

raw_event
 → schema validation
 → semantic enrichment
 → ontology mapping
 → graph ingestion

2026 production design updates

Expected enrichment outputs include:

Graph-ready extraction contract

Use a graph-native response envelope so Gemini output can be written directly into Firestore-backed knowledge graph structures:

{
  "entities": [
    {
      "id": "customer_123",
      "type": "Customer",
      "attributes": {
        "email_hash": "sha256:...",
        "journey_stage": "consideration"
      }
    }
  ],
  "relationships": [
    {
      "source": "customer_123",
      "target": "event_456",
      "type": "EXPERIENCED",
      "timestamp": "2026-03-01T10:00:00Z",
      "sequence": 4
    }
  ]
}

This pattern eliminates parser glue code and makes Gemini a first-class ETL layer.

4) Identity Resolution Graph

Identity entities

Core relationships

Resolution strategy

  1. Deterministic matching: email hash, CRM ID, authenticated logins.
  2. Probabilistic matching: IP/device/behavior similarity.
  3. Cluster and merge into persistent customer nodes with confidence scores.
  4. Assign stable canonical IDs before journey materialization.

Critical implementation rule

Do not let journeys materialize on unresolved identities. Entity resolution quality matters more than extraction richness because fragmented identities corrupt downstream attribution, personalization, and graph traversal.

5) Customer Journey Profile Model

Represent journeys as ordered touchpoint sequences:

Ad Impression
→ Ad Click
→ Landing Page
→ Product View
→ Email Signup
→ Email Click
→ Purchase

Profile structure:

Temporal graph requirement

Customer journeys must be modeled as time-indexed graphs, not flat event lists. Every relationship that represents a touchpoint should carry:

This enables:

Journey profile must be materialized as a KG node

The customer journey profile is not just an application-side summary. It should exist as a first-class graph node with stable identity and explicit relationships.

Minimum journey node

{
  "id": "journey_customer_123",
  "type": "Journey",
  "attributes": {
    "customer_id": "customer_123",
    "journey_stage": "consideration",
    "touchpoint_ids": ["event_001", "event_002", "event_003"],
    "sequence_depth": 3,
    "started_at": "2026-03-01T10:00:00Z",
    "updated_at": "2026-03-10T12:00:00Z"
  }
}

Minimum journey relationships

6) Ontology Design Baseline

Core ontology entities

Core relationships

Schema lifecycle management

The ontology should be treated as a living contract:

This is especially important when new channels, support artifacts, or document-derived entities are added into the intake layer.

7) Firestore Graph Materialization Pattern

Firestore should be treated as a denormalized operational graph store.

Core collections

/entities/{id}
/relationships/{id}

Adjacency indexes

/entities/{id}/edges_out
/entities/{id}/edges_in

Materialized views

/customers/{id}/journey_summary
/customers/{id}/state

Implementation guidelines

KG upload order for journey-aware writes

When uploading Gemini Pipeline v2 output into the knowledge graph, write in this order:

  1. Customer, Journey, Touchpoint, Campaign, and Channel nodes.
  2. Customer -> Journey ownership edges.
  3. Journey -> Touchpoint sequence membership edges.
  4. Touchpoint -> Touchpoint temporal PRECEDES edges.
  5. Attribution and channel edges such as Touchpoint -> Campaign and Campaign -> Channel.

This ensures that the customer journey profile is already in place when downstream agents query the graph.

8) Provenance and Governance Requirements

Attach provenance metadata to every node and edge write:

Minimum governance controls:

9) Event-Driven Runtime Pattern

Recommended production runtime:

New data
  → Pub/Sub / Webhook intake
  → Normalization worker
  → Gemini extraction worker
  → Entity resolution + graph builder
  → Firestore writes
  → Journey/materialized-view updater

Operational requirements:

Concrete KG upsert payload

The graph builder should emit both journey profiles and explicit node/relationship write plans:

{
  "journeys": [
    {
      "id": "journey_customer_123",
      "customer_id": "customer_123",
      "stage": "consideration",
      "touchpoint_ids": ["event_001", "event_002", "event_003"]
    }
  ],
  "entities": [
    { "id": "customer_123", "type": "Customer" },
    { "id": "journey_customer_123", "type": "Journey" },
    { "id": "event_003", "type": "Touchpoint" }
  ],
  "relationships": [
    { "source": "customer_123", "target": "journey_customer_123", "type": "HAS_JOURNEY" },
    { "source": "journey_customer_123", "target": "event_003", "type": "HAS_TOUCHPOINT", "sequence": 3 },
    { "source": "event_002", "target": "event_003", "type": "PRECEDES", "sequence": 3 }
  ]
}

10) Attribution on Top of the Graph

Build multi-touch attribution with weighted graph traversals:

Purchase
↑
Email Click
↑
Meta Click
↑
Google Search
↑
Organic Visit

Weighting function baseline:

edge_weight = engagement_score × recency × channel_influence

Supports:

9) Platform Mapping for This Repository

Ingestion and connectors

Unified event schema and validation

Identity and journey processing

KG build and semantic enrichment

Decisioning and control plane

10) Implementation Rules (Non-Negotiable)

  1. Separate entities from metrics using dimension/fact modeling.
  2. Store raw payloads indefinitely to enable replay and model retraining.
  3. Version event schemas (event_schema_v1, event_schema_v2, ...).
  4. Maintain a master campaign registry for cross-platform canonical IDs.
  5. Use streaming ingestion by default with near-real-time identity + graph updates.

Strategic Direction

Mature marketing intelligence stacks converge on:

event streams + identity graph + semantic ontology + AI decision layer

This enables autonomous optimization, real-time journey orchestration, explainable attribution, and robust cross-channel decision support.

References

  1. Meta Marketing API Insights best practices: https://developers.facebook.com/docs/marketing-api/insights/best-practices/
  2. Meta Ads API update summary (June 2025): https://windsor.ai/documentation/facebook-ads-meta-api-updates-june-10-2025/
  3. Conversions API real-time discussion: https://www.adamigo.ai/blog/ultimate-guide-real-time-data-meta-ads-ai
  4. Google Ads API batch processing best practices: https://developers.google.com/google-ads/api/docs/batch-processing/best-practices
  5. Google Ads data strength best practices: https://support.google.com/google-ads/answer/16517525
  6. Google Cloud Enterprise Knowledge Graph overview: https://docs.cloud.google.com/enterprise-knowledge-graph/docs/overview
← All docsView source on GitHub →