Transformation & Processing → Stream Processing

Identity Resolution (Stitching)

Links events from different devices, sessions, and channels back to the same real customer — the mechanism that makes 'unified profile' actually true.

High-Level Design

Identity Resolution is what makes cross-channel personalization possible.

Data Source
Sessionization
Hands off session-scoped events
→
Ingestion
Identity & Profile Service
Owns the identity graph queried here
→
Processing
Identity Resolution
.NET Core Stream Worker — identity graph stitching
→
Foundation
Data Quality Checks
Runs next, validating the stitched event
→
Intelligence
Azure Database for PostgreSQL
Identity graph storage
→
Activation
Unified Customer Profile
What makes cross-channel personalization possible

💼 Business Context

  • Without stitching, the same customer looks like several different anonymous strangers across web, mobile, and in-store — the core promise of the platform depends on this working
  • Directly determines the accuracy of every downstream personalization, LTV, and segmentation calculation
  • Owned by the Identity & Profile Service team

🔌 Technical Overview

Stitching runs as a Stream Worker step that resolves an incoming event's anonymous_id or device identifier against the identity graph maintained by the Identity & Profile Service (backed by Azure Database for PostgreSQL). A deterministic match (e.g., login, loyalty scan, email click) creates or strengthens an edge between an anonymous ID and a known user_id; a probabilistic match (shared device/IP heuristics) is scored but flagged with lower confidence and never silently treated as certain.

Match Types

Deterministic (login, loyalty ID, email) Probabilistic (device/IP heuristics, scored)

💾 Identity Graph Edge

{
  "anonymous_id": "a1b2c3d4",
  "user_id": "cust_004821",
  "match_type": "deterministic",
  "match_source": "login",
  "confidence": 1.0,
  "linked_at": "2026-08-01T10:20:01Z"
}

🔗 Integration Points

  • Identity & Profile Service (.NET Core API) — owns the identity graph
  • Azure Database for PostgreSQL — identity graph storage
  • .NET Core Stream Worker — calls the Identity API inline during processing to resolve/attach identity
  • Every downstream consumer (Analytics, AI & Insights, Activation) — reads the resolved user_id, not raw anonymous IDs

🧰 Services Consumed

  • Owning microservice — Cxos.Processing.Api (see the Full Application Service Map)
  • Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)

⚠️ Non-Functional Considerations

  • Scale: identity graph lookups are cached for active sessions to avoid a database round-trip per event
  • Latency: adds a small, cache-mitigated latency cost per event in exchange for correctness
  • Reliability: an identity-service outage degrades to anonymous-only processing rather than blocking the pipeline
  • Security/Privacy: probabilistic matches are never used for regulated purposes (e.g., legal consent decisions) — only deterministic matches carry that weight

🎯 Enterprise Example

A customer browses on mobile, then completes checkout on desktop while logged in. Deterministic stitching via the login event links both anonymous IDs to the same user_id, so the Analytics API correctly attributes the purchase to the mobile browsing session that originally drove it — instead of showing two unrelated visitors.

← Back to Stream Processing