Marketing Journey Knowledge Graph Architecture (2026)
Purpose
This document integrates the latest marketing data pipeline patterns into the MIZOKI stack so raw ad/email/web signals are transformed into a governed, ontology-driven customer journey knowledge graph.
Target Reference Architecture
Marketing APIs / Webhooks
↓
Connector Layer
↓
Event Stream
↓
Schema Normalization
↓
Identity Resolution Engine
↓
Gemini Semantic Structuring
↓
Ontology Mapping
↓
Knowledge Graph
↓
AI Decision Agents
1) Source Extraction Patterns
Meta / Facebook Marketing API
Recommended ingestion mode: async + streaming hybrid.
- Trigger asynchronous Insights report jobs.
- Poll report status and fetch paginated result sets.
- Stream normalized records to the event bus.
POST /act_{account_id}/insights
fields=impressions,clicks,spend,actions
time_increment=1
level=ad
Implementation guidelines
- Use daily backfills plus hourly deltas.
- Use incremental filters (
updated_time) for mutation-safe re-ingestion. - Pull metadata (campaign/adset/ad objects) independently from facts.
- Pin attribution window settings to avoid metric drift after 2025 reach/attribution updates.
- Add Conversions API events into the same stream for low-latency (15–30 minute) optimization loops.
Google Ads / GMP APIs
Query standard: GAQL with fact/dimension separation.
Fact tables:
ad_performancekeyword_performanceconversion_events
Dimension tables:
campaignad_groupadcreativegeodevice
Implementation guidelines
- Apply related operations as atomic batches.
- Preserve ordering of batch requests to prevent transient inconsistent states.
Email engagement APIs
Open events should be treated as weak signals. Prioritize:
- click events
- landing events
- scroll depth
- session duration
Reference flow:
ESP webhook → event queue → engagement classifier
Programmatic / DSP logs
Prefer log-level exports and retain raw event streams before rollups.
Common event types:
bid_requestimpressionviewable_impressionclickvideo_quartileconversion
Web / app tracking
Prefer server-side tracking with stable identifiers:
event_idevent_timestampanonymous_iduser_iddevice_idsession_idreferrercampaigncontent_id
2) Event Normalization Contract
All inbound events should be transformed into the canonical schema before enrichment and graph writes.
event_id
event_type
event_timestamp
actor_id
session_id
source_system
channel
campaign_id
creative_id
properties
value
Example:
event_type: ad_click
channel: google_ads
campaign_id: 812
actor_id: anon_73
landing_url: /product/shoes
3) Gemini Semantic Structuring Layer
Apply LLM-assisted enrichment after schema validation:
raw_event
→ schema validation
→ semantic enrichment
→ ontology mapping
→ graph ingestion
2026 production design updates
- Treat Gemini structured output as a deterministic transformation contract, not an unstructured reasoning step.
- Use a single JSON schema across:
- extraction
- service-to-service exchange
- Firestore entity and relationship writes
- Prefer schema features that support evolving ontologies:
$refanyOf- strict field ordering
- For document-heavy sources, use Gemini as both:
- document parser
- semantic extractor
- Separate extraction into two passes when quality matters:
1.
pass_1_extract— entities, events, relationships 2.pass_2_graph_refine— duplicate resolution, edge refinement, ontology correction
Expected enrichment outputs include:
- entity class (Campaign, Creative, Audience, Transaction)
- inferred intent (Acquisition, Retention, Reactivation)
- audience taxonomy labels
- creative/media type
- normalized channel mapping
Graph-ready extraction contract
Use a graph-native response envelope so Gemini output can be written directly into Firestore-backed knowledge graph structures:
{
"entities": [
{
"id": "customer_123",
"type": "Customer",
"attributes": {
"email_hash": "sha256:...",
"journey_stage": "consideration"
}
}
],
"relationships": [
{
"source": "customer_123",
"target": "event_456",
"type": "EXPERIENCED",
"timestamp": "2026-03-01T10:00:00Z",
"sequence": 4
}
]
}
This pattern eliminates parser glue code and makes Gemini a first-class ETL layer.
4) Identity Resolution Graph
Identity entities
CustomerEmailDeviceCookieAdPlatformUserCRMContact
Core relationships
Customer -[:HAS_EMAIL]-> EmailCustomer -[:USES_DEVICE]-> DeviceDevice -[:LINKED_COOKIE]-> CookieCustomer -[:ASSOCIATED_WITH]-> AdPlatformUser
Resolution strategy
- Deterministic matching: email hash, CRM ID, authenticated logins.
- Probabilistic matching: IP/device/behavior similarity.
- Cluster and merge into persistent customer nodes with confidence scores.
- Assign stable canonical IDs before journey materialization.
Critical implementation rule
Do not let journeys materialize on unresolved identities. Entity resolution quality matters more than extraction richness because fragmented identities corrupt downstream attribution, personalization, and graph traversal.
5) Customer Journey Profile Model
Represent journeys as ordered touchpoint sequences:
Ad Impression
→ Ad Click
→ Landing Page
→ Product View
→ Email Signup
→ Email Click
→ Purchase
Profile structure:
customer_idjourney_idjourney_stagetouchpoints[]engagement_scoreconversion_probabilitylifetime_value_estimate
Temporal graph requirement
Customer journeys must be modeled as time-indexed graphs, not flat event lists. Every relationship that represents a touchpoint should carry:
timestampsequence- optional session/window context
This enables:
- funnel analysis
- path prediction
- behavioral clustering
- graph-enriched retrieval
Journey profile must be materialized as a KG node
The customer journey profile is not just an application-side summary. It should exist as a first-class graph node with stable identity and explicit relationships.
Minimum journey node
{
"id": "journey_customer_123",
"type": "Journey",
"attributes": {
"customer_id": "customer_123",
"journey_stage": "consideration",
"touchpoint_ids": ["event_001", "event_002", "event_003"],
"sequence_depth": 3,
"started_at": "2026-03-01T10:00:00Z",
"updated_at": "2026-03-10T12:00:00Z"
}
}
Minimum journey relationships
Customer -[:HAS_JOURNEY]-> JourneyJourney -[:HAS_TOUCHPOINT]-> TouchpointTouchpoint -[:PRECEDES]-> TouchpointCustomer -[:EXPERIENCED]-> TouchpointTouchpoint -[:ATTRIBUTED_TO]-> CampaignCampaign -[:RUNS_ON]-> Channel
6) Ontology Design Baseline
Core ontology entities
CustomerAccountCampaignCreativeChannelEventTouchpointSessionProductTransaction
Core relationships
Customer -[:EXPERIENCED]-> EventEvent -[:BELONGS_TO]-> CampaignCampaign -[:RUNS_ON]-> ChannelCustomer -[:PURCHASED]-> Product
Schema lifecycle management
The ontology should be treated as a living contract:
- version schemas explicitly
- annotate schema changes
- reprocess historical records when new entity/relationship fields are introduced
- reindex graph materializations after ontology evolution
This is especially important when new channels, support artifacts, or document-derived entities are added into the intake layer.
7) Firestore Graph Materialization Pattern
Firestore should be treated as a denormalized operational graph store.
Core collections
/entities/{id}
/relationships/{id}
Adjacency indexes
/entities/{id}/edges_out
/entities/{id}/edges_in
Materialized views
/customers/{id}/journey_summary
/customers/{id}/state
Implementation guidelines
- Upsert canonical entities before writing relationships.
- Maintain bidirectional adjacency indexes for traversal without joins.
- Materialize customer journey summaries for hot-path reads.
- Denormalize aggressively; do not depend on runtime joins for decisioning paths.
- Store provenance on both nodes and edges.
KG upload order for journey-aware writes
When uploading Gemini Pipeline v2 output into the knowledge graph, write in this order:
Customer,Journey,Touchpoint,Campaign, andChannelnodes.Customer -> Journeyownership edges.Journey -> Touchpointsequence membership edges.Touchpoint -> TouchpointtemporalPRECEDESedges.- Attribution and channel edges such as
Touchpoint -> CampaignandCampaign -> Channel.
This ensures that the customer journey profile is already in place when downstream agents query the graph.
8) Provenance and Governance Requirements
Attach provenance metadata to every node and edge write:
source_systemapi_endpointingestion_timestampapi_versionconfidence_scoredata_lineageprivacy_scope
Minimum governance controls:
- retention policy enforcement
- PII tagging and classification
- GDPR/CCPA deletion workflows
- immutable audit history
9) Event-Driven Runtime Pattern
Recommended production runtime:
New data
→ Pub/Sub / Webhook intake
→ Normalization worker
→ Gemini extraction worker
→ Entity resolution + graph builder
→ Firestore writes
→ Journey/materialized-view updater
Operational requirements:
- retry queues for transient extraction failures
- dead-letter queues for schema violations
- versioned schemas for backward compatibility
- chunking + context stitching for long documents
- running entity registry to improve cross-chunk consistency
Concrete KG upsert payload
The graph builder should emit both journey profiles and explicit node/relationship write plans:
{
"journeys": [
{
"id": "journey_customer_123",
"customer_id": "customer_123",
"stage": "consideration",
"touchpoint_ids": ["event_001", "event_002", "event_003"]
}
],
"entities": [
{ "id": "customer_123", "type": "Customer" },
{ "id": "journey_customer_123", "type": "Journey" },
{ "id": "event_003", "type": "Touchpoint" }
],
"relationships": [
{ "source": "customer_123", "target": "journey_customer_123", "type": "HAS_JOURNEY" },
{ "source": "journey_customer_123", "target": "event_003", "type": "HAS_TOUCHPOINT", "sequence": 3 },
{ "source": "event_002", "target": "event_003", "type": "PRECEDES", "sequence": 3 }
]
}
10) Attribution on Top of the Graph
Build multi-touch attribution with weighted graph traversals:
Purchase
↑
Email Click
↑
Meta Click
↑
Google Search
↑
Organic Visit
Weighting function baseline:
edge_weight = engagement_score × recency × channel_influence
Supports:
- multi-touch attribution
- causal inference inputs
- budget/media optimization
9) Platform Mapping for This Repository
Ingestion and connectors
- Extend source adapters in
src/cells/cell02/andsrc/cells/cell02/integrations/for async Meta pulls, GAQL extraction, webhook ingestion, and log-level DSP connectors.
Unified event schema and validation
- Keep schema contracts in
templates/schemas/and enforce viasrc/moa/workflows/schema_validate_workflow.py.
Identity and journey processing
- Use
src/cells/cell02/enhanced_identity_resolver.pyfor deterministic + probabilistic merge logic. - Use
src/cells/cell02/customer_journey_profiler.pyfor sequence construction and journey scoring.
KG build and semantic enrichment
- Use
services/gemini-kg-pipeline/src/gemini/extraction_engine.pyandservices/gemini-kg-pipeline/src/kg/populator.pyto apply semantic structuring and graph writes. - Keep ontology-aware schema definitions in
services/gemini-kg-pipeline/src/schemas/.
Decisioning and control plane
- Feed graph outputs into
services/marketing-control-plane/for attribution-aware orchestration and optimization loops.
10) Implementation Rules (Non-Negotiable)
- Separate entities from metrics using dimension/fact modeling.
- Store raw payloads indefinitely to enable replay and model retraining.
- Version event schemas (
event_schema_v1,event_schema_v2, ...). - Maintain a master campaign registry for cross-platform canonical IDs.
- Use streaming ingestion by default with near-real-time identity + graph updates.
Strategic Direction
Mature marketing intelligence stacks converge on:
event streams + identity graph + semantic ontology + AI decision layer
This enables autonomous optimization, real-time journey orchestration, explainable attribution, and robust cross-channel decision support.
References
- Meta Marketing API Insights best practices: https://developers.facebook.com/docs/marketing-api/insights/best-practices/
- Meta Ads API update summary (June 2025): https://windsor.ai/documentation/facebook-ads-meta-api-updates-june-10-2025/
- Conversions API real-time discussion: https://www.adamigo.ai/blog/ultimate-guide-real-time-data-meta-ads-ai
- Google Ads API batch processing best practices: https://developers.google.com/google-ads/api/docs/batch-processing/best-practices
- Google Ads data strength best practices: https://support.google.com/google-ads/answer/16517525
- Google Cloud Enterprise Knowledge Graph overview: https://docs.cloud.google.com/enterprise-knowledge-graph/docs/overview