Catalog
The searchable index of every table, its location, schema, and owner across the entire lakehouse — the platform's map of itself.
High-Level Design
The Catalog is how anyone finds out what data already exists.
💼 Business Context
- Without a catalog, "does this data already exist somewhere" becomes tribal knowledge — the catalog makes discovery self-service
- Reduces duplicate data pipelines built because a team didn't know an existing table already had what they needed
- Owned by Data Governance / Platform Engineering
🔌 Technical Overview
The catalog is implemented on Azure Purview, which automatically scans and registers every table in the Data Lakehouse — location, schema, partition spec, and owner metadata. The Query & Analytics Engine resolves logical table names to physical Iceberg table locations through the catalog rather than hardcoded paths, and the developer portal exposes catalog search so any engineer can discover existing tables before building a new pipeline.
Catalog Contents
💾 Catalog Entry
{
"table": "curated.identity_stitched_events",
"owner_team": "platform-engineering",
"location": "abfss://lakehouse@cxosdata.dfs.core.windows.net/curated/events",
"row_count_estimate": 2847193021,
"last_updated": "2026-08-01T09:00:00Z"
}
🔗 Integration Points
- Azure Purview — the catalog implementation
- Query & Analytics Engine — resolves table names via catalog lookups
- Developer portal — surfaces catalog search to engineers
- Data Lakehouse — every table is auto-registered on creation
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Azure Database for PostgreSQL (policy/retention state)
⚠️ Non-Functional Considerations
- Scale: catalog scanning runs incrementally, not a full rescan of the lakehouse on every change
- Latency: catalog lookups are cached and fast; not a bottleneck for query planning
- Reliability: catalog staleness is bounded by the scan schedule — typically near-real-time for new tables
- Security/Privacy: catalog entries expose schema and metadata, not row-level data, so browsing the catalog is lower-risk than querying a table directly
🎯 Enterprise Example
A new analyst wants revenue data and searches the catalog for 'revenue' instead of asking around — finding the existing marts.revenue_by_category_daily table and its owning team in seconds, avoiding a duplicate pipeline.