Unified Data Foundation → CXOS Data Lakehouse

Partitioning & Clustering

The physical data-layout strategy that makes queries against a multi-billion-row table fast instead of a full scan.

High-Level Design

Partitioning is the difference between a fast dashboard and a timeout.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Transformation & Processing
Writes data using the partition strategy
→
Foundation
Partitioning & Clustering
Iceberg hidden partitioning by date/customer
→
Intelligence
Query & Analytics Engine
Partition pruning speeds up filtered queries
→
Activation
Interactive Dashboard Latency
The direct beneficiary of good partitioning

💼 Business Context

  • The difference between a dashboard that loads in a second and one that times out often comes down entirely to partitioning strategy
  • Keeps query costs (compute time) manageable as the lakehouse grows into the billions of rows
  • Owned by Data Engineering

🔌 Technical Overview

Tables are partitioned primarily by event date (supporting the common 'last N days' query pattern) and, for high-cardinality access patterns, clustered by customer key. Iceberg's hidden partitioning means queries don't need to explicitly reference partition columns — the engine automatically prunes irrelevant partitions based on filter predicates in the query, unlike older Hive-style partitioning that required exact partition-column matches.

Strategy

Partition by event date Cluster by customer_key Automatic partition pruning

💾 Partition Specification

{
  "table": "curated.identity_stitched_events",
  "partition-spec": [{ "field": "event_date", "transform": "day" }],
  "sort-order": [{ "field": "customer_key" }]
}

🔗 Integration Points

  • Apache Iceberg — hidden partitioning implementation
  • Query & Analytics Engine — partition pruning happens transparently during query planning
  • Rollups & Aggregations — batch jobs are designed around the date-partition boundary
  • Azure Data Lake Storage Gen2 — physical file layout follows the partition structure

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Infrastructure (see the Full Application Service Map)
  • Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself

⚠️ Non-Functional Considerations

  • Scale: good partitioning is what allows query performance to stay roughly constant as total table size grows
  • Latency: a well-pruned query touches only the relevant partitions, not the entire table history
  • Reliability: partition strategy is set at table-design time and reviewed like any other schema decision
  • Security/Privacy: overly fine-grained partitioning can create many small files, an operational cost addressed by Lifecycle Management

🎯 Enterprise Example

A dashboard querying 'last 7 days of revenue' against a 3-billion-row events table returns in under a second, because partition pruning means the query engine only ever touches 7 days' worth of files instead of the full table history.

← Back to CXOS Data Lakehouse