Partitioning & Clustering
The physical data-layout strategy that makes queries against a multi-billion-row table fast instead of a full scan.
High-Level Design
Partitioning is the difference between a fast dashboard and a timeout.
💼 Business Context
- The difference between a dashboard that loads in a second and one that times out often comes down entirely to partitioning strategy
- Keeps query costs (compute time) manageable as the lakehouse grows into the billions of rows
- Owned by Data Engineering
🔌 Technical Overview
Tables are partitioned primarily by event date (supporting the common 'last N days' query pattern) and, for high-cardinality access patterns, clustered by customer key. Iceberg's hidden partitioning means queries don't need to explicitly reference partition columns — the engine automatically prunes irrelevant partitions based on filter predicates in the query, unlike older Hive-style partitioning that required exact partition-column matches.
Strategy
💾 Partition Specification
{
"table": "curated.identity_stitched_events",
"partition-spec": [{ "field": "event_date", "transform": "day" }],
"sort-order": [{ "field": "customer_key" }]
}
🔗 Integration Points
- Apache Iceberg — hidden partitioning implementation
- Query & Analytics Engine — partition pruning happens transparently during query planning
- Rollups & Aggregations — batch jobs are designed around the date-partition boundary
- Azure Data Lake Storage Gen2 — physical file layout follows the partition structure
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Infrastructure(see the Full Application Service Map) - Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself
⚠️ Non-Functional Considerations
- Scale: good partitioning is what allows query performance to stay roughly constant as total table size grows
- Latency: a well-pruned query touches only the relevant partitions, not the entire table history
- Reliability: partition strategy is set at table-design time and reviewed like any other schema decision
- Security/Privacy: overly fine-grained partitioning can create many small files, an operational cost addressed by Lifecycle Management
🎯 Enterprise Example
A dashboard querying 'last 7 days of revenue' against a 3-billion-row events table returns in under a second, because partition pruning means the query engine only ever touches 7 days' worth of files instead of the full table history.