Anomaly Detection
Continuously monitors key metrics and data-quality signals for statistically unusual behavior and raises it before a human would otherwise notice.
High-Level Design
Anomaly Detection watches the marts layer so a bad number gets caught before a dashboard is.
💼 Business Context
- Catches metric and data-quality regressions — a broken connector, a mis-tagged event, a real business anomaly — before a stakeholder spots it in a dashboard and asks an unplanned question
- Reduces mean-time-to-detect for revenue or volume anomalies from 'whenever someone notices' to minutes
- Owned by Analytics Engineering / Data Science, alert routing owned by Operational Services
🔌 Technical Overview
An Azure Machine Learning-hosted model (a seasonal-decomposition and z-score ensemble, retrained periodically) scores every registered Semantic Layer metric on each dbt refresh, comparing the latest value against its expected range given historical seasonality. The scoring job runs as a .NET Core-orchestrated Azure Container Apps Job — a Docker container invoking the Azure ML scoring endpoint — and any metric outside its confidence band is written to an anomalies table that Alerts & Notifications polls.
Detected Anomaly Types
💾 Anomaly Record
{
"metric": "net_revenue",
"dimension": { "region": "APAC" },
"expected_range": [42000, 58000],
"observed": 12400,
"severity": "high",
"detected_at": "2026-08-02T07:15:00Z"
}
🔗 Integration Points
- Rollups & Aggregations — supplies the metric time series being scored
- Azure Machine Learning — hosts and serves the anomaly-scoring model
- Alerts & Notifications (Operational Services) — routes detected anomalies to the owning team
- Data Quality Monitoring (Operational Services) — a related but distinct signal source; anomalies here are metric-level, not row-level
🧰 Services Consumed
- Owning microservice —
Cxos.Intelligence.Api(see the Full Application Service Map) - Database — Azure Data Explorer/Kusto (scoring time series) + Azure Database for PostgreSQL
⚠️ Non-Functional Considerations
- Scale: scoring runs once per metric per refresh cycle, so cost scales with catalog size, not raw event volume
- Latency: anomalies are detected within one refresh cycle of the underlying metric — typically within the hour
- Reliability: the model is retrained on a rolling window so seasonal patterns (e.g., holiday spikes) do not get permanently flagged as anomalies
- Security/Privacy: scoring operates on aggregated metrics, not row-level customer data, keeping the model's inputs low-sensitivity
🎯 Enterprise Example
An APAC connector silently breaks overnight, dropping reported revenue to a quarter of its expected range. Anomaly Detection flags it within the hour and Alerts & Notifications pages the on-call data engineer — instead of the drop being discovered two days later when Finance asks why the weekly number looks wrong.