Unified Data Foundation → CXOS Data Lakehouse

Zero-Copy Architecture

The principle that data is stored once and queried by many engines/consumers, instead of being copied into each tool that needs it.

High-Level Design

One physical copy, many query engines — no per-tool duplication.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Transformation & Processing
Writes once to the lakehouse
→
Foundation
Zero-Copy Architecture
One physical copy, many query engines
→
Intelligence
Query & Analytics Engine
Reads the same physical files
→
Activation
Every Consumer
AI & Insights, BI tools, ad hoc analysis — no duplication

💼 Business Context

  • Eliminates the storage cost and staleness risk of maintaining separate copies of data for each tool (a warehouse copy, a BI-tool extract, an ML-training copy)
  • Every consumer sees the same data at the same freshness — no "which copy is right" debates
  • Owned by Platform Architecture

🔌 Technical Overview

Because every zone is stored as Iceberg tables on Azure Data Lake Storage Gen2, any engine that speaks the Iceberg format reads the same physical Parquet files directly — there's no ETL step that copies data into a separate proprietary warehouse format. The Query & Analytics Engine, ad hoc Spark jobs, Snowflake (via Iceberg external tables), and Azure Machine Learning training pipelines all read from the same underlying storage, differing only in compute engine, not data location.

Benefits

No duplicate storage cost No cross-copy staleness Single access-control surface

💾 Multi-Engine Access

-- Query & Analytics Engine (DataFusion)
SELECT * FROM curated.identity_stitched_events;

-- Ad hoc Spark job, same physical table
spark.read.format("iceberg").load("curated.identity_stitched_events")

🔗 Integration Points

  • Azure Data Lake Storage Gen2 — the single physical storage layer
  • Apache Iceberg — makes multi-engine access possible without copies
  • Query & Analytics Engine, Azure Machine Learning, ad hoc tools — all read the same tables
  • Governance & Security — single point of access control since there is only one copy to secure

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Infrastructure (see the Full Application Service Map)
  • Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself

⚠️ Non-Functional Considerations

  • Scale: avoids the storage cost multiplication of N tools each holding their own copy
  • Latency: no ETL delay between "data is ready" and "every tool can see it" — it is the same moment
  • Reliability: eliminates an entire class of bugs caused by copies drifting out of sync
  • Security/Privacy: a single access-control surface is easier to audit correctly than N separate copies with N separate permission systems

🎯 Enterprise Example

Previously, the BI team maintained their own extract, and the ML team maintained a separate training dataset — both drifting from the source over time. With zero-copy architecture, both read the same Iceberg tables directly, and a data quality fix in the curated zone is instantly visible to both teams simultaneously.

← Back to CXOS Data Lakehouse