Unified Data Foundation → CXOS Data Lakehouse

Raw → Curated → Analytics Data Marts

The three-tier structure every event passes through inside the lakehouse — raw as received, curated after cleaning and modeling, and marts for fast analytical consumption.

High-Level Design

Every zone lives on the same storage, differing only in transformation stage.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Transformation & Processing
Cleaned, enriched, modeled events
→
Foundation
Lakehouse Zones
Raw → Curated → Marts on Azure Data Lake Storage Gen2
→
Intelligence
Query & Analytics Engine
Reads primarily from the marts zone
→
Activation
Every Destination
Ultimately sourced from lakehouse data

💼 Business Context

  • Separating raw/curated/marts lets the platform keep an unmodified historical record while still serving fast, clean data to BI tools
  • Protects against a common failure mode: 'fixing' data in place and losing the ability to reprocess from the original source
  • Owned by Data Engineering

🔌 Technical Overview

All three zones live on Azure Data Lake Storage Gen2 as Iceberg tables, differing only in transformation stage, not storage technology. The raw zone is an append-only, schema-on-write copy of validated events (written directly by the Ingestion API). The curated zone holds deduplicated, identity-stitched, quality-checked data produced by Transformation & Processing. The marts zone holds the pre-aggregated rollups and dimensional models described under Batch Processing, optimized for the Query & Analytics Engine's access patterns.

Zones

Raw (append-only) Curated (modeled, deduped) Marts (pre-aggregated)

💾 Zone Table Naming

raw.events_2026_08
curated.identity_stitched_events
marts.revenue_by_category_daily

🔗 Integration Points

  • Azure Data Lake Storage Gen2 — physical storage for all three zones
  • Ingestion API — writes directly to the raw zone
  • Transformation & Processing (Stream/Batch Workers) — populate curated and marts zones
  • Query & Analytics Engine — primarily reads from marts, falls back to curated for ad hoc queries

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Infrastructure (see the Full Application Service Map)
  • Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself

⚠️ Non-Functional Considerations

  • Scale: each zone scales independently; raw grows unboundedly while marts stays compact via aggregation
  • Latency: raw zone is near-real-time; curated lags by the Stream Worker's processing time; marts lags by the batch schedule
  • Reliability: because raw is immutable and complete, curated/marts can always be rebuilt from it if a bug is found
  • Security/Privacy: access policies differ by zone — raw often has the tightest access since it is least processed/redacted

🎯 Enterprise Example

A data science team needs a metric that doesn't exist in any mart yet. Because the curated zone retains full event-level detail, they can compute it directly rather than waiting for a new ETL pipeline to be built and backfilled.

← Back to CXOS Data Lakehouse