Unified Data Foundation → CXOS Data Lakehouse

Open Table Format (Iceberg)

The open, vendor-neutral table format underlying every zone of the lakehouse — what makes 'zero-copy' and 'query engine of your choice' actually possible.

High-Level Design

Iceberg is the metadata layer that gives table semantics to files in blob storage.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Transformation & Processing
Writes tables in Iceberg format
→
Foundation
Open Table Format (Iceberg)
Table metadata layer over Azure Data Lake Storage Gen2 Parquet files
→
Intelligence
Query & Analytics Engine
DataFusion reads Iceberg tables directly
→
Activation
Any Iceberg-compatible Engine
Spark, Trino, Snowflake can all attach

💼 Business Context

  • Avoids vendor lock-in to a single proprietary warehouse format — the same physical data can be queried by multiple engines
  • Enables the platform's zero-copy architecture: no separate copy needed per consuming tool
  • Owned by Data Engineering / Platform Architecture

🔌 Technical Overview

Apache Iceberg provides a metadata layer (manifest files, snapshots) on top of Parquet files in Azure Data Lake Storage Gen2, giving SQL-table semantics (schema, partitions, ACID transactions) to files in blob storage. The Query & Analytics Engine's DataFusion runtime reads Iceberg tables natively; the format's open specification also lets Spark, Trino, or Snowflake attach to the same tables without a data copy or export step.

Iceberg Capabilities

ACID transactions Hidden partitioning Snapshot isolation Engine-agnostic metadata Snowflake external tables

💾 Table Metadata Reference

{
  "table": "curated.identity_stitched_events",
  "format-version": 2,
  "current-snapshot-id": 5821093741,
  "location": "abfss://lakehouse@cxosdata.dfs.core.windows.net/curated/events"
}

🔗 Integration Points

  • Azure Data Lake Storage Gen2 — underlying Parquet file storage
  • Apache Iceberg — the table format/metadata specification
  • Query & Analytics Engine (DataFusion) — primary read engine
  • Any Iceberg-compatible engine (Spark, Trino, Snowflake via external tables) — can attach without migration

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Infrastructure (see the Full Application Service Map)
  • Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself

⚠️ Non-Functional Considerations

  • Scale: Iceberg's manifest-based metadata scales to very large tables without the small-file listing problems of raw Parquet/Hive tables
  • Latency: metadata operations (snapshot listing, partition pruning) are fast regardless of table size
  • Reliability: ACID transactions prevent readers from seeing partially written data during concurrent writes
  • Security/Privacy: table-level access control is enforced by the Governance & Security layer on top of the format

🎯 Enterprise Example

A data science team wants to run a one-off Spark job against curated event data. Because it's stored as Iceberg, they attach Spark directly to the existing tables — no export, no separate copy, no waiting on a data engineering ticket.

← Back to CXOS Data Lakehouse