Raw → Curated → Analytics Data Marts
The three-tier structure every event passes through inside the lakehouse — raw as received, curated after cleaning and modeling, and marts for fast analytical consumption.
High-Level Design
Every zone lives on the same storage, differing only in transformation stage.
💼 Business Context
- Separating raw/curated/marts lets the platform keep an unmodified historical record while still serving fast, clean data to BI tools
- Protects against a common failure mode: 'fixing' data in place and losing the ability to reprocess from the original source
- Owned by Data Engineering
🔌 Technical Overview
All three zones live on Azure Data Lake Storage Gen2 as Iceberg tables, differing only in transformation stage, not storage technology. The raw zone is an append-only, schema-on-write copy of validated events (written directly by the Ingestion API). The curated zone holds deduplicated, identity-stitched, quality-checked data produced by Transformation & Processing. The marts zone holds the pre-aggregated rollups and dimensional models described under Batch Processing, optimized for the Query & Analytics Engine's access patterns.
Zones
💾 Zone Table Naming
raw.events_2026_08 curated.identity_stitched_events marts.revenue_by_category_daily
🔗 Integration Points
- Azure Data Lake Storage Gen2 — physical storage for all three zones
- Ingestion API — writes directly to the raw zone
- Transformation & Processing (Stream/Batch Workers) — populate curated and marts zones
- Query & Analytics Engine — primarily reads from marts, falls back to curated for ad hoc queries
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Infrastructure(see the Full Application Service Map) - Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself
⚠️ Non-Functional Considerations
- Scale: each zone scales independently; raw grows unboundedly while marts stays compact via aggregation
- Latency: raw zone is near-real-time; curated lags by the Stream Worker's processing time; marts lags by the batch schedule
- Reliability: because raw is immutable and complete, curated/marts can always be rebuilt from it if a bug is found
- Security/Privacy: access policies differ by zone — raw often has the tightest access since it is least processed/redacted
🎯 Enterprise Example
A data science team needs a metric that doesn't exist in any mart yet. Because the curated zone retains full event-level detail, they can compute it directly rather than waiting for a new ETL pipeline to be built and backfilled.