Event Collection
The front door of the platform — the endpoint that receives every event from every SDK, connector, and webhook before anything else happens to it.
High-Level Design
Event Collection is the acknowledgment boundary between callers and the platform.
💼 Business Context
- Every downstream capability depends on collection being rock-solid — an outage here is an outage for the entire platform
- The single place where "did we receive this event at all" can be answered, which matters for data-completeness audits
- Owned by Platform Engineering; treated as a tier-1 production service
🔌 Technical Overview
Event Collection is the ASP.NET Core Web API endpoint fronted by Azure API Management, built and shipped as a Docker image so every environment (local dev, staging, production AKS nodes) runs the identical container. It accepts events over HTTPS (SDKs), gRPC (high-throughput server callers), and via the webhook receivers described elsewhere. Every request is authenticated, given a receipt timestamp and a server-assigned event_id if the caller didn't supply one, and acknowledged with a 202 Accepted before any validation or enrichment happens — collection is deliberately decoupled from processing so a downstream slowdown never blocks intake. The Application Insights SDK is wired into the endpoint out of the box, so every request gets a distributed-tracing correlation ID that follows it through validation, Event Hubs, and every downstream .NET Core service — the same trace ID that shows up in an exception report if something later fails.
Supported Transports
💾 Collection Acknowledgment
HTTP/1.1 202 Accepted
{
"event_id": "9f2c1e6a-2b41-4e9d-8f3a-1c7d9a0b6e2f",
"received_at": "2026-08-01T10:22:14.203Z",
"status": "queued"
}
🔗 Integration Points
- Azure API Management — ingress, auth, throttling for all transports
- .NET Core Ingestion API — Docker container on AKS, the collection endpoint itself
- Azure Event Hubs — immediate hand-off after acknowledgment
- Application Insights — per-request distributed tracing and exception telemetry
- Azure Monitor — collection-tier health and throughput dashboards
🧰 Services Consumed
- Owning microservice —
Cxos.Ingestion.Application(see the Full Application Service Map) - Database — Azure Cache for Redis (cache only, no system-of-record database)
⚠️ Non-Functional Considerations
- Scale: AKS horizontal pod autoscaling keeps collection latency flat under load; acknowledgment happens before any heavy processing
- Latency: target is sub-100ms acknowledgment at the 99th percentile
- Reliability: collection is stateless and horizontally scaled — any pod can handle any request
- Security/Privacy: authentication happens before any payload parsing to avoid processing unauthenticated traffic
🎯 Enterprise Example
During a product launch, event volume spikes 15x. Because collection only authenticates and hands off to Azure Event Hubs — without waiting for downstream processing — acknowledgment latency stays flat even while the Stream Worker queue temporarily backs up, so no client-side SDK ever times out or drops events.