Data Quality Monitoring
Continuously validates row-level and schema-level data quality — nulls, referential integrity, dbt test failures — surfacing issues before they reach a dashboard or model.
High-Level Design
Data Quality Monitoring catches row-level problems; Anomaly Detection catches metric-level ones.
💼 Business Context
- Catches broken pipelines and bad data at the row level before they propagate into a metric, a model feature, or a customer-facing decision
- Gives data owners a scorecard for their tables instead of quality being invisible until someone complains
- Owned by Data Engineering, scorecards reviewed per domain team
🔌 Technical Overview
Every dbt model in Transformation & Processing carries dbt test assertions (not-null, uniqueness, referential integrity, accepted-value ranges), executed as part of the same scheduled run that materializes the model. A .NET Core service (Docker container on Azure Container Apps) aggregates test results into a per-table quality scorecard and publishes failures to Azure Service Bus for Alerts & Notifications, distinguishing this row/schema-level signal from Anomaly Detection's metric-level scoring.
Test Types
💾 Quality Scorecard Entry
{
"table": "curated.identity_stitched_events",
"tests_run": 24,
"tests_passed": 23,
"failure": { "test": "not_null:customer_key", "failed_rows": 142 },
"run_at": "2026-08-02T03:00:00Z"
}
🔗 Integration Points
- dbt test — the assertion framework underlying every check
- Rollups & Aggregations / Data Modeling — the scheduled runs quality checks execute alongside
- Alerts & Notifications — receives test-failure events for routing
- Catalog — quality scorecard is surfaced alongside each table's catalog entry
🧰 Services Consumed
- Owning microservice —
Cxos.Operations.Api(see the Full Application Service Map) - Database — Azure Database for PostgreSQL + Azure Cosmos DB Table API + Azure Data Explorer
⚠️ Non-Functional Considerations
- Scale: test execution overhead scales with model count and test count, tuned to run within the existing batch window
- Latency: quality issues are caught within one dbt run cycle of the data landing — typically within the hour for high-frequency models
- Reliability: a critical test failure can be configured to block downstream models from running on bad data (dbt's failure-propagation), preventing a bad batch from silently cascading
- Security/Privacy: quality scorecards report aggregate pass/fail counts, not the offending row-level data itself, in team-visible summaries
🎯 Enterprise Example
A connector starts sending customer_key as null for 5% of a table's new rows after an upstream change. The not-null test fails on the next dbt run, the model is blocked from feeding downstream marts, and the owning team is paged — instead of a silently degraded profile match rate being discovered weeks later.