Unified Data Foundation → Metadata Layer

Catalog

The searchable index of every table, its location, schema, and owner across the entire lakehouse — the platform's map of itself.

High-Level Design

The Catalog is how anyone finds out what data already exists.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Data Lakehouse
Every table is registered here automatically
→
Foundation
Catalog
Azure Purview — searchable table/schema index
→
Intelligence
Query & Analytics Engine
Resolves table locations via the catalog
→
Activation
Developer Portal
Self-service table discovery

💼 Business Context

  • Without a catalog, "does this data already exist somewhere" becomes tribal knowledge — the catalog makes discovery self-service
  • Reduces duplicate data pipelines built because a team didn't know an existing table already had what they needed
  • Owned by Data Governance / Platform Engineering

🔌 Technical Overview

The catalog is implemented on Azure Purview, which automatically scans and registers every table in the Data Lakehouse — location, schema, partition spec, and owner metadata. The Query & Analytics Engine resolves logical table names to physical Iceberg table locations through the catalog rather than hardcoded paths, and the developer portal exposes catalog search so any engineer can discover existing tables before building a new pipeline.

Catalog Contents

Table location & schema Owner metadata Partition spec Last-updated timestamp

💾 Catalog Entry

{
  "table": "curated.identity_stitched_events",
  "owner_team": "platform-engineering",
  "location": "abfss://lakehouse@cxosdata.dfs.core.windows.net/curated/events",
  "row_count_estimate": 2847193021,
  "last_updated": "2026-08-01T09:00:00Z"
}

🔗 Integration Points

  • Azure Purview — the catalog implementation
  • Query & Analytics Engine — resolves table names via catalog lookups
  • Developer portal — surfaces catalog search to engineers
  • Data Lakehouse — every table is auto-registered on creation

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Api (see the Full Application Service Map)
  • Database — ADLS Gen2 (Iceberg) + Azure Database for PostgreSQL (policy/retention state)

⚠️ Non-Functional Considerations

  • Scale: catalog scanning runs incrementally, not a full rescan of the lakehouse on every change
  • Latency: catalog lookups are cached and fast; not a bottleneck for query planning
  • Reliability: catalog staleness is bounded by the scan schedule — typically near-real-time for new tables
  • Security/Privacy: catalog entries expose schema and metadata, not row-level data, so browsing the catalog is lower-risk than querying a table directly

🎯 Enterprise Example

A new analyst wants revenue data and searches the catalog for 'revenue' instead of asking around — finding the existing marts.revenue_by_category_daily table and its owning team in seconds, avoiding a duplicate pipeline.

← Back to Metadata Layer