Identity Graph
The graph of every known identifier for a customer — device IDs, cookies, emails, loyalty IDs — and the resolved links between them that make "one customer" possible.
High-Level Design
Identity Graph is the resolution layer the Unified Customer Profile is keyed on.
💼 Business Context
- Without a maintained identity graph, "unified" profile is a fiction — this is the mechanism that actually links an anonymous web cookie to a known loyalty member
- Improves attribution accuracy and prevents the same person being double-counted as two customers in reporting
- Owned by Data Engineering / Customer Data Platform team
🔌 Technical Overview
Identity edges (device_id↔email, cookie↔loyalty_id, etc.) produced by the Identity Resolution step during Transformation & Processing are persisted as a versioned graph table in the curated zone. The Analytics/AI API — a Docker container running on Azure Container Apps — exposes a graph-traversal endpoint that resolves any known identifier to its canonical customer_key, using Azure Database for PostgreSQL's recursive CTE support for multi-hop traversal and Azure Cache for Redis to cache hot lookups.
Identifier Types
💾 Identity Resolution Query
GET /identity/resolve?device_id=dev_88f2
{
"customer_key": "cust_004821",
"resolved_via": ["device_id", "email_hash"],
"confidence": "high",
"linked_identifiers": 4
}
🔗 Integration Points
- Data Sources' Identity Resolution — produces the raw edges this graph consumes
- curated.identity_edges — the Iceberg table storing graph edges
- Unified Customer Profile — resolves customer_key via this graph before every profile read
- AI & Insights — anomaly detection flags identity edges with unusually low confidence
🧰 Services Consumed
- Owning microservice —
Cxos.Profile.Api(see the Full Application Service Map) - Database — Azure Cosmos DB (Core API + Gremlin API) + Azure Cache for Redis
⚠️ Non-Functional Considerations
- Scale: edge count grows faster than customer count (many identifiers per customer), so the graph table is partitioned by customer_key for traversal performance
- Latency: single-hop resolution is cache-served in single-digit milliseconds; rare multi-hop traversals fall back to PostgreSQL and complete in under 100ms
- Reliability: low-confidence merges are flagged rather than auto-applied, preventing one bad match from silently merging two different customers
- Security/Privacy: raw identifiers (email, phone) are stored hashed, and the graph itself is classified PII under Data Classification
🎯 Enterprise Example
A customer browses anonymously on mobile, then logs in on desktop hours later. The identity graph links the anonymous device_id session to the newly authenticated email, so their browsing behavior — not just the desktop session — informs the propensity score the AI & Insights engine computes minutes later.