Databricks / Redshift
Serves data-science and legacy-warehouse teams standardized on Databricks or Redshift, via the same open-format access pattern as any other Iceberg-compatible engine.
High-Level Design
Databricks attaches like Spark; Redshift receives a scheduled load like BigQuery.
💼 Business Context
- Data science teams standardized on Databricks get direct Spark access to the same tables everything else reads, without a duplicate ML-specific copy
- Legacy Redshift-based BI teams remain supported during a migration window without blocking their existing dashboards
- Owned by Data Engineering / Data Science Platform
🔌 Technical Overview
Databricks connects to the lakehouse the same way any Iceberg-compatible Spark engine does — direct attachment to Azure Data Lake Storage Gen2 tables, no export job required, matching the pattern already established for ad hoc Spark access under Zero-Copy Architecture. Redshift, without native Iceberg support, receives a scheduled load job (a Docker container on Azure Container Apps Jobs) that reads via the Query & Analytics Engine and writes into Redshift-native tables, used primarily to support legacy dashboards during migration to the lakehouse-native path.
Delivery Modes
💾 Databricks Attachment Example
# Databricks notebook, direct Iceberg read — no CXOS export job involved
df = spark.read.format("iceberg") \
.load("abfss://lakehouse@cxosdata.dfs.core.windows.net/marts/order_fact")
🔗 Integration Points
- Zero-Copy Architecture — the same principle extended to Databricks' Spark engine
- Query & Analytics Engine — read path for the Redshift scheduled load
- Governance & Security — grants and reviews Databricks workspace access to lakehouse tables
- Predictive Models — Databricks is a common environment for ad hoc feature exploration ahead of formal dbt Python model changes
🧰 Services Consumed
- Owning microservice —
Cxos.Connectors.BatchExport(see the Full Application Service Map) - No dedicated database — stateless connector (see Platform Connectors above)
⚠️ Non-Functional Considerations
- Scale: Databricks attachment has no export-size ceiling; Redshift loads are sized like any scheduled batch export
- Latency: Databricks reflects the lakehouse near-real-time; Redshift lags by its load schedule
- Reliability: the Databricks path has no separate copy to go stale; the Redshift path is monitored like any other batch export job
- Security/Privacy: Databricks workspace access is governed and audited the same as any direct lakehouse consumer, not treated as an external export
🎯 Enterprise Example
A data science team exploring a new feature for the next Predictive Models iteration attaches Databricks directly to curated tables for ad hoc exploration, while the BI team's existing Redshift dashboards keep working unchanged off a nightly load job until their planned migration to querying the lakehouse directly.