Ingestion Layer → Edge Network

Queue & Retry

Absorbs bursts and outages between the edge and downstream processing, so a slow or failing consumer never becomes data loss.

High-Level Design

Queue & Retry is the buffer that decouples ingestion from every consumer.

Data Source
Consent Enforcement
Hands off a policy-cleared event
→
Ingestion
Queue & Retry
Azure Event Hubs buffering + Polly retries
→
Processing
.NET Core Stream Worker
Primary consumer, reading from the buffer
→
Foundation
Data Lakehouse
Ultimate durable destination
→
Intelligence
Consumer Offset Tracking
Safe backlog catch-up after an outage
→
Activation
Zero Data-Loss Guarantee
What makes the platform's reliability SLA possible

💼 Business Context

  • Guarantees that a downstream incident (a bad deploy, a database outage) doesn't translate into lost customer data
  • Removes the pressure to over-provision every downstream service for worst-case traffic, since the queue absorbs the spike
  • Owned by Platform Engineering; its throughput/retention settings are a direct input to the platform's data-loss SLA

🔌 Technical Overview

Once an event passes validation and enrichment, it's published to Azure Event Hubs, which retains events for a configurable window (default 7 days) independent of consumer health. Consumers (the Stream Worker, batch jobs) track their own read offset, so a consumer outage simply means a backlog to catch up on, not lost data. Client-side (Web/Mobile SDK) and server-side (Cxos.Ingestion.Client) callers additionally retry failed sends with exponential backoff via Polly before an event is ever considered lost.

Retry Layers

Client-side SDK retry Server SDK Polly policy Event Hubs retention buffer Dead-letter fallback

💾 Retry Policy Configuration

services.AddHttpClient()
  .AddPolicyHandler(Policy
    .Handle()
    .WaitAndRetryAsync(5, attempt =>
       TimeSpan.FromSeconds(Math.Pow(2, attempt))));

🔗 Integration Points

  • Azure Event Hubs — the durable buffer between ingestion and every consumer
  • Polly retry/circuit-breaker policies — built into Cxos.Ingestion.Client
  • Consumer offset tracking (Stream Worker, batch jobs) — enables safe backlog catch-up
  • Dead-letter topic — final fallback for events that exhaust retries

🧰 Services Consumed

  • Owning microservice — Cxos.Ingestion.Application (see the Full Application Service Map)
  • Database — Azure Cache for Redis (cache only, no system-of-record database)

⚠️ Non-Functional Considerations

  • Scale: Event Hubs throughput units scale independently of consumer capacity, so ingestion never has to slow down to match a slower downstream stage
  • Latency: under normal conditions, queue dwell time is milliseconds; during an incident it becomes backlog depth, not data loss
  • Reliability: this is the mechanism that makes the rest of the platform's reliability claims possible
  • Security/Privacy: queued events remain subject to the same encryption-at-rest and access policies as stored data, not a lower-security holding area

🎯 Enterprise Example

A bad deploy takes the Stream Worker offline for 40 minutes. Because Event Hubs retains events for 7 days, there is zero data loss — the worker resumes from its last committed offset and catches up on the backlog within 10 minutes of redeployment, with no customer-visible impact beyond a short delay in real-time personalization.

← Back to Edge Network