Unified Data Foundation → Metadata Layer

Lineage

Traces exactly which upstream sources and transformations produced any given table or field — the platform's answer to "where did this number come from".

High-Level Design

Lineage turns "where did this come from" into a queryable graph.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Transformation & Processing
Every job reports its inputs/outputs
→
Foundation
Lineage
Azure Purview — traces tables back to their sources
→
Intelligence
Impact Analysis
What breaks if this table changes
→
Activation
Root-Cause Debugging
Where a bad number is traced back to its origin

💼 Business Context

  • When a number looks wrong, lineage is what lets someone trace it back to its source instead of guessing
  • Enables impact analysis before a schema or pipeline change: "what downstream tables/dashboards does this affect"
  • Owned by Data Governance / Platform Engineering

🔌 Technical Overview

Every batch and stream job registers its inputs and outputs with Azure Purview's lineage API as part of its execution, building a directed graph from raw source tables through every transformation to final marts and dashboards. This is automatic instrumentation built into the shared batch/stream job frameworks, not something each team has to remember to add manually.

Lineage Graph Nodes

Source tables Transformation jobs Output tables Downstream dashboards

💾 Lineage Query Result

marts.revenue_by_category_daily
  <- job: rollups_aggregations (daily)
  <- curated.order_fact
    <- job: data_modeling (hourly)
    <- curated.identity_stitched_events
      <- job: stream_worker (real-time)
      <- raw.events

🔗 Integration Points

  • Azure Purview — lineage graph storage and API
  • Batch/Stream job frameworks — auto-instrumented to report lineage on every run
  • Catalog — lineage is displayed alongside catalog entries
  • Impact-analysis tooling — used before schema changes or deprecations

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Api (see the Full Application Service Map)
  • Database — ADLS Gen2 (Iceberg) + Azure Database for PostgreSQL (policy/retention state)

⚠️ Non-Functional Considerations

  • Scale: lineage graph grows with the number of tables and jobs, not event volume — manageable at any realistic pipeline count
  • Latency: lineage updates as jobs run; current to the pipeline execution level, not the event level
  • Reliability: auto-instrumentation means lineage doesn't silently go stale because a team forgot to document it
  • Security/Privacy: lineage metadata is generally lower-sensitivity than the data itself, but access is still governed

🎯 Enterprise Example

A dashboard shows an unexpected revenue drop. Instead of guessing, the analyst traces the metric's lineage back through the marts, curated, and raw tables and finds a specific upstream connector had a data-quality issue that day — cutting root-cause time from hours to minutes.

← Back to Metadata Layer