Alerts & Notifications
The routing layer that turns an anomaly, a workflow state change, or a data-quality failure into a notification the right person or team actually sees.
High-Level Design
Alerts & Notifications is the last-mile delivery for every operational signal generated elsewhere in the platform.
💼 Business Context
- A detected anomaly or data-quality issue has no value if it sits in a table nobody looks at — this is what makes signals actionable
- Reduces mean-time-to-detect and mean-time-to-resolve by routing directly to the owning team rather than a shared, ignorable channel
- Owned by Platform Engineering / SRE
🔌 Technical Overview
A .NET Core routing service (Docker container on Azure Container Apps) subscribes to Azure Service Bus topics published by Anomaly Detection, Data Quality Monitoring, and the Workflow Engine, matches each event against ownership metadata from the Catalog to determine the responsible team, and fans out to the configured channel (email via SendGrid, Slack webhook, or PagerDuty for severity-critical alerts) with configurable per-team severity thresholds to avoid alert fatigue.
Delivery Channels
💾 Alert Routing Rule
{
"source": "anomaly_detection",
"severity_threshold": "high",
"owner_team": "platform-engineering",
"channels": ["slack", "pagerduty"],
"dedup_window_minutes": 30
}
🔗 Integration Points
- Azure Service Bus — the event backbone alerts are published and consumed through
- Anomaly Detection, Data Quality Monitoring, Workflow Engine — primary upstream signal sources
- Catalog (Metadata Layer) — owner metadata used to route an alert to the correct team
- SendGrid, Slack, PagerDuty — external delivery integrations
🧰 Services Consumed
- Owning microservice —
Cxos.Operations.Api(see the Full Application Service Map) - Database — Azure Database for PostgreSQL + Azure Cosmos DB Table API + Azure Data Explorer
⚠️ Non-Functional Considerations
- Scale: alert volume is bounded by monitored-metric and workflow count, not raw event volume
- Latency: alerts are delivered within seconds to minutes of the triggering condition, well inside the mean-time-to-detect targets this system exists to hit
- Reliability: a deduplication window prevents the same underlying issue from paging a team repeatedly within a short window
- Security/Privacy: alert payloads summarize the issue (metric name, severity) rather than embedding raw customer data, keeping notification channels low-sensitivity
🎯 Enterprise Example
A critical data-quality failure is detected at 2am. Alerts & Notifications pages the owning team's on-call engineer via PagerDuty within a minute, rather than the issue being discovered the next morning when someone happens to check a dashboard.