Transformation & Processing → Batch Processing

Backfills

Reprocesses historical data when a new source is onboarded, a bug is fixed, or a business rule changes — without waiting for months of new data to accumulate.

High-Level Design

Backfills bring history in line with the current logic, safely.

Data Source
Data Lakehouse (raw zone)
Source of truth for reprocessing
→
Ingestion
dbt Core
Same model SQL used for backfill and production runs
→
Processing
Backfills
dbt run with a historical date-range variable
→
Foundation
Iceberg Time Travel / Versioning
Enables safe, validated cutover
→
Intelligence
Same Transformation Logic
Reused, not reimplemented, for consistency
→
Activation
Corrected Historical Reporting
Brings history in line with current logic

💼 Business Context

  • Makes a newly onboarded data source immediately useful with full history, instead of starting from a blank slate
  • Lets a bug fix or business-logic change apply retroactively, keeping historical reporting consistent with current definitions
  • Owned by Data Engineering, requested by whichever team needs the correction or onboarding

🔌 Technical Overview

Backfills use dbt's variable-driven date-range pattern: dbt run --select model_name --vars '{"start_date": ..., "end_date": ...}' reprocesses a historical window using the exact same model SQL as production — there's no separate backfill codebase that can drift out of sync with the real pipeline. The job runs as a Docker container on Azure Container Apps Jobs, writing to a new Iceberg snapshot; because the lakehouse's Iceberg tables support time travel, the previous snapshot remains queryable and the 'current' pointer only switches over once dbt test confirms the backfilled output against a validation target.

Common Triggers

New source onboarding Bug fix in derivation logic Business rule change Schema evolution

💾 dbt Backfill Command

dbt run --select identity_stitched_events \
  --vars '{"start_date": "2026-01-01", "end_date": "2026-07-31"}' \
  --target validation

🔗 Integration Points

  • dbt Core — the same model SQL used for backfill and production runs, parameterized by date range
  • Iceberg time travel / versioning — enables safe, validated cutover before the new snapshot becomes current
  • Docker container on Azure Container Apps Jobs — long-running backfill execution
  • dbt test — runs against the validation target before the backfilled data is promoted

🧰 Services Consumed

  • Owning microservice — Cxos.Processing.Api (see the Full Application Service Map)
  • Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)

⚠️ Non-Functional Considerations

  • Scale: large backfills (months to years of history) are chunked by date range and run as parallel dbt invocations
  • Latency: backfills are explicitly not real-time — they're scheduled, monitored, long-running jobs
  • Reliability: table versioning allows validation before cutover and instant rollback if a backfill produces unexpected results
  • Security/Privacy: backfills follow the same access-control and audit-logging requirements as any other write to the lakehouse

🎯 Enterprise Example

A fix to the identity-stitching confidence threshold is deployed. Rather than accepting seven months of under-stitched historical profiles, dbt run reprocesses that window with the corrected model against a validation target, dbt test confirms the output, and the snapshot is promoted — bringing historical data in line with the corrected logic without a multi-month wait or a second codebase to maintain.

← Back to Batch Processing