Backfills
Reprocesses historical data when a new source is onboarded, a bug is fixed, or a business rule changes — without waiting for months of new data to accumulate.
High-Level Design
Backfills bring history in line with the current logic, safely.
💼 Business Context
- Makes a newly onboarded data source immediately useful with full history, instead of starting from a blank slate
- Lets a bug fix or business-logic change apply retroactively, keeping historical reporting consistent with current definitions
- Owned by Data Engineering, requested by whichever team needs the correction or onboarding
🔌 Technical Overview
Backfills use dbt's variable-driven date-range pattern: dbt run --select model_name --vars '{"start_date": ..., "end_date": ...}' reprocesses a historical window using the exact same model SQL as production — there's no separate backfill codebase that can drift out of sync with the real pipeline. The job runs as a Docker container on Azure Container Apps Jobs, writing to a new Iceberg snapshot; because the lakehouse's Iceberg tables support time travel, the previous snapshot remains queryable and the 'current' pointer only switches over once dbt test confirms the backfilled output against a validation target.
Common Triggers
💾 dbt Backfill Command
dbt run --select identity_stitched_events \
--vars '{"start_date": "2026-01-01", "end_date": "2026-07-31"}' \
--target validation
🔗 Integration Points
- dbt Core — the same model SQL used for backfill and production runs, parameterized by date range
- Iceberg time travel / versioning — enables safe, validated cutover before the new snapshot becomes current
- Docker container on Azure Container Apps Jobs — long-running backfill execution
- dbt test — runs against the validation target before the backfilled data is promoted
🧰 Services Consumed
- Owning microservice —
Cxos.Processing.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)
⚠️ Non-Functional Considerations
- Scale: large backfills (months to years of history) are chunked by date range and run as parallel dbt invocations
- Latency: backfills are explicitly not real-time — they're scheduled, monitored, long-running jobs
- Reliability: table versioning allows validation before cutover and instant rollback if a backfill produces unexpected results
- Security/Privacy: backfills follow the same access-control and audit-logging requirements as any other write to the lakehouse
🎯 Enterprise Example
A fix to the identity-stitching confidence threshold is deployed. Rather than accepting seven months of under-stitched historical profiles, dbt run reprocesses that window with the corrected model against a validation target, dbt test confirms the output, and the snapshot is promoted — bringing historical data in line with the corrected logic without a multi-month wait or a second codebase to maintain.