Sessionization
Groups a customer's raw events into logical sessions — the unit most analytics, personalization, and journey logic actually operate on.
High-Level Design
Sessionization turns a raw event stream into the unit analytics actually uses.
💼 Business Context
- Almost every meaningful metric (bounce rate, session duration, funnel conversion) is defined at the session level, not the raw event level
- Removes the burden from every downstream team to reimplement session-boundary logic themselves
- Owned by Platform Engineering, in consultation with Analytics
🔌 Technical Overview
The Stream Worker groups events sharing the same identity (user_id or anonymous_id) and channel into a session using a configurable inactivity timeout (default 30 minutes) — a standard sliding-window session model. A session_start derived event is emitted on the first event of a new session and session_end when the timeout fires, both written back onto Azure Event Hubs so downstream consumers (including the Analytics API) can react to session boundaries directly.
Session Rules
💾 Derived Session Event
{
"event": "session_end",
"session_id": "sess_77af",
"user_id": "cust_004821",
"started_at": "2026-08-01T10:15:02Z",
"ended_at": "2026-08-01T10:41:19Z",
"event_count": 14
}
🔗 Integration Points
- .NET Core Stream Worker — where sessionization runs, stateful windowing over Azure Event Hubs
- Azure Cache for Redis — holds in-flight session state until the inactivity timeout fires
- Analytics API — consumes session_start/session_end as first-class events
- Identity & Profile Service — sessions roll up into the customer's broader activity timeline
🧰 Services Consumed
- Owning microservice —
Cxos.Processing.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)
⚠️ Non-Functional Considerations
- Scale: session state is sharded by customer ID, scaling horizontally with the Stream Worker's partition count
- Latency: session_end fires within the configured timeout window, not instantly — an inherent tradeoff of inactivity-based sessionization
- Reliability: session state survives a worker restart via checkpointed Redis state, not in-memory only
- Security/Privacy: session identifiers are opaque and rotate per session — never a persistent customer identifier on their own
🎯 Enterprise Example
A customer browses for 20 minutes, leaves the tab open unattended for an hour, then returns and completes a purchase. Sessionization correctly treats this as two separate sessions, so the Analytics API attributes the purchase to a fresh, high-intent return session rather than diluting a single artificially long one.