Lineage
Traces exactly which upstream sources and transformations produced any given table or field — the platform's answer to "where did this number come from".
High-Level Design
Lineage turns "where did this come from" into a queryable graph.
💼 Business Context
- When a number looks wrong, lineage is what lets someone trace it back to its source instead of guessing
- Enables impact analysis before a schema or pipeline change: "what downstream tables/dashboards does this affect"
- Owned by Data Governance / Platform Engineering
🔌 Technical Overview
Every batch and stream job registers its inputs and outputs with Azure Purview's lineage API as part of its execution, building a directed graph from raw source tables through every transformation to final marts and dashboards. This is automatic instrumentation built into the shared batch/stream job frameworks, not something each team has to remember to add manually.
Lineage Graph Nodes
💾 Lineage Query Result
marts.revenue_by_category_daily
<- job: rollups_aggregations (daily)
<- curated.order_fact
<- job: data_modeling (hourly)
<- curated.identity_stitched_events
<- job: stream_worker (real-time)
<- raw.events
🔗 Integration Points
- Azure Purview — lineage graph storage and API
- Batch/Stream job frameworks — auto-instrumented to report lineage on every run
- Catalog — lineage is displayed alongside catalog entries
- Impact-analysis tooling — used before schema changes or deprecations
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Azure Database for PostgreSQL (policy/retention state)
⚠️ Non-Functional Considerations
- Scale: lineage graph grows with the number of tables and jobs, not event volume — manageable at any realistic pipeline count
- Latency: lineage updates as jobs run; current to the pipeline execution level, not the event level
- Reliability: auto-instrumentation means lineage doesn't silently go stale because a team forgot to document it
- Security/Privacy: lineage metadata is generally lower-sensitivity than the data itself, but access is still governed
🎯 Enterprise Example
A dashboard shows an unexpected revenue drop. Instead of guessing, the analyst traces the metric's lineage back through the marts, curated, and raw tables and finds a specific upstream connector had a data-quality issue that day — cutting root-cause time from hours to minutes.