Unified Data Foundation → CXOS Data Lakehouse

Schema Evolution

The ability to add, rename, or reorder columns in a lakehouse table without rewriting existing data or breaking existing queries.

High-Level Design

Schema Evolution lets tables change shape as the business does.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Event Schema & Registry
Governs what changes are allowed
→
Foundation
Schema Evolution
Iceberg's in-place schema change support
→
Intelligence
Query & Analytics Engine
Old and new queries both keep working
→
Activation
Uninterrupted Downstream Consumers
No forced migration window

💼 Business Context

  • The business changes constantly, and the lakehouse's table structure has to change with it without a disruptive migration
  • Removes a major source of "can we add this field" friction between business teams and data engineering
  • Owned by Data Engineering

🔌 Technical Overview

Iceberg tracks column changes (add, drop, rename, reorder, widen type) as metadata operations rather than rewriting existing Parquet files. Old files remain readable under the new schema (missing new columns read as null), and this capability is what the Event Schema & Registry's compatibility rules are built to safely take advantage of at the lakehouse layer, not just the event-contract layer.

Supported Changes

Add column Rename column Reorder columns Widen type (int → long)

💾 Schema Evolution Operation

ALTER TABLE curated.orders
ADD COLUMN tax_amount DECIMAL(10,2);
-- Existing files remain valid; new column reads as NULL for old rows

🔗 Integration Points

  • Apache Iceberg — provides the underlying schema evolution mechanism
  • Event Schema & Registry — governs which changes are safe to propagate into the lakehouse
  • Data Modeling — dimensional tables evolve using this capability as business entities change
  • Query & Analytics Engine — must handle nulls gracefully for newly added columns on historical rows

🧰 Services Consumed

  • Owning microservice — Cxos.Foundation.Infrastructure (see the Full Application Service Map)
  • Database — Azure Data Lake Storage Gen2 (Iceberg) — the lakehouse itself

⚠️ Non-Functional Considerations

  • Scale: schema changes are metadata-only operations, independent of table size — no multi-hour rewrite of historical data
  • Latency: a schema change takes effect immediately for new writes; historical data is reinterpreted, not rewritten
  • Reliability: schema evolution is additive-safe by design — existing queries don't break when a column is added
  • Security/Privacy: newly added columns must still go through the same PII classification process as any new field

🎯 Enterprise Example

The Commerce team needs to add tax_amount to the orders table. Schema evolution lets this happen instantly without rewriting years of historical order data or breaking any dashboard currently querying the table.

← Back to CXOS Data Lakehouse