Data Quality Checks
Continuous, automated checks on event data as it streams through — catching quality regressions before they reach a dashboard or model.
High-Level Design
Data Quality Checks run in parallel with processing, watching for regressions.
💼 Business Context
- Silent data quality regressions (a broken tracking call, a schema drift) erode trust in the platform faster than almost any other failure mode
- Lets teams catch their own integration bugs within minutes instead of discovering them in a monthly business review
- Owned by Platform Engineering, with checks configurable per event type by the owning domain team
🔌 Technical Overview
Data Quality Checks run as a parallel consumer on the same Event Hubs stream, applying both rule-based checks (null rates, value-range checks, referential checks like 'does this product_id exist in the catalog') and statistical checks (volume anomaly detection comparing current throughput to a rolling baseline). Failures don't block the pipeline — they raise an Azure Monitor alert and populate a data-quality dashboard the owning team can act on.
Check Types
💾 Data Quality Alert
{
"check": "null_rate",
"field": "properties.product_id",
"event_type": "product_viewed",
"threshold": 0.02,
"observed": 0.18,
"status": "breached",
"window": "last_15_minutes"
}
🔗 Integration Points
- Parallel consumer on Azure Event Hubs, independent of the main processing path
- Azure Monitor — alerting when a check breaches its threshold
- Data-quality dashboard (built on the Analytics API) — visibility for domain teams
- Event Schema & Registry — check rules are partly derived from the registered schema's constraints
🧰 Services Consumed
- Owning microservice —
Cxos.Processing.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)
⚠️ Non-Functional Considerations
- Scale: runs as a lightweight parallel stream, sized independently of the primary processing path
- Latency: checks operate on rolling windows (typically 15 minutes), trading a small detection delay for statistical stability
- Reliability: a data-quality check failure never blocks or slows the primary pipeline — it's purely observational
- Security/Privacy: checks operate on field presence/shape/statistics, not on raw PII values
🎯 Enterprise Example
A marketing team's tag-manager update accidentally stops populating product_id on 18% of product_viewed events. The null-rate check breaches its 2% threshold within 15 minutes and alerts the team directly, instead of the gap being discovered a month later when a report looks strange.