Open Table Format (Iceberg)
The open, vendor-neutral table format underlying every zone of the lakehouse — what makes 'zero-copy' and 'query engine of your choice' actually possible.
High-Level Design
Iceberg is the metadata layer that gives table semantics to files in blob storage.
💼 Business Context
- Avoids vendor lock-in to a single proprietary warehouse format — the same physical data can be queried by multiple engines
- Enables the platform's zero-copy architecture: no separate copy needed per consuming tool
- Owned by Data Engineering / Platform Architecture
🔌 Technical Overview
Apache Iceberg provides a metadata layer (manifest files, snapshots) on top of Parquet files in Azure Data Lake Storage Gen2, giving SQL-table semantics (schema, partitions, ACID transactions) to files in blob storage. The Query & Analytics Engine's DataFusion runtime reads Iceberg tables natively; the format's open specification also lets Spark, Trino, or Snowflake attach to the same tables without a data copy or export step.
Iceberg Capabilities
💾 Table Metadata Reference
{
"table": "curated.identity_stitched_events",
"format-version": 2,
"current-snapshot-id": 5821093741,
"location": "abfss://lakehouse@cxosdata.dfs.core.windows.net/curated/events"
}
🔗 Integration Points
- Azure Data Lake Storage Gen2 — underlying Parquet file storage
- Apache Iceberg — the table format/metadata specification
- Query & Analytics Engine (DataFusion) — primary read engine
- Any Iceberg-compatible engine (Spark, Trino, Snowflake via external tables) — can attach without migration
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Infrastructure(see the Full Application Service Map) - Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself
⚠️ Non-Functional Considerations
- Scale: Iceberg's manifest-based metadata scales to very large tables without the small-file listing problems of raw Parquet/Hive tables
- Latency: metadata operations (snapshot listing, partition pruning) are fast regardless of table size
- Reliability: ACID transactions prevent readers from seeing partially written data during concurrent writes
- Security/Privacy: table-level access control is enforced by the Governance & Security layer on top of the format
🎯 Enterprise Example
A data science team wants to run a one-off Spark job against curated event data. Because it's stored as Iceberg, they attach Spark directly to the existing tables — no export, no separate copy, no waiting on a data engineering ticket.