Queue & Retry
Absorbs bursts and outages between the edge and downstream processing, so a slow or failing consumer never becomes data loss.
High-Level Design
Queue & Retry is the buffer that decouples ingestion from every consumer.
💼 Business Context
- Guarantees that a downstream incident (a bad deploy, a database outage) doesn't translate into lost customer data
- Removes the pressure to over-provision every downstream service for worst-case traffic, since the queue absorbs the spike
- Owned by Platform Engineering; its throughput/retention settings are a direct input to the platform's data-loss SLA
🔌 Technical Overview
Once an event passes validation and enrichment, it's published to Azure Event Hubs, which retains events for a configurable window (default 7 days) independent of consumer health. Consumers (the Stream Worker, batch jobs) track their own read offset, so a consumer outage simply means a backlog to catch up on, not lost data. Client-side (Web/Mobile SDK) and server-side (Cxos.Ingestion.Client) callers additionally retry failed sends with exponential backoff via Polly before an event is ever considered lost.
Retry Layers
💾 Retry Policy Configuration
services.AddHttpClient() .AddPolicyHandler(Policy .Handle () .WaitAndRetryAsync(5, attempt => TimeSpan.FromSeconds(Math.Pow(2, attempt))));
🔗 Integration Points
- Azure Event Hubs — the durable buffer between ingestion and every consumer
- Polly retry/circuit-breaker policies — built into Cxos.Ingestion.Client
- Consumer offset tracking (Stream Worker, batch jobs) — enables safe backlog catch-up
- Dead-letter topic — final fallback for events that exhaust retries
🧰 Services Consumed
- Owning microservice —
Cxos.Ingestion.Application(see the Full Application Service Map) - Database — Azure Cache for Redis (cache only, no system-of-record database)
⚠️ Non-Functional Considerations
- Scale: Event Hubs throughput units scale independently of consumer capacity, so ingestion never has to slow down to match a slower downstream stage
- Latency: under normal conditions, queue dwell time is milliseconds; during an incident it becomes backlog depth, not data loss
- Reliability: this is the mechanism that makes the rest of the platform's reliability claims possible
- Security/Privacy: queued events remain subject to the same encryption-at-rest and access policies as stored data, not a lower-security holding area
🎯 Enterprise Example
A bad deploy takes the Stream Worker offline for 40 minutes. Because Event Hubs retains events for 7 days, there is zero data loss — the worker resumes from its last committed offset and catches up on the backlog within 10 minutes of redeployment, with no customer-visible impact beyond a short delay in real-time personalization.