Destinations & Activation → Batch / File Exports

Databricks / Redshift

Serves data-science and legacy-warehouse teams standardized on Databricks or Redshift, via the same open-format access pattern as any other Iceberg-compatible engine.

High-Level Design

Databricks attaches like Spark; Redshift receives a scheduled load like BigQuery.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Rollups & Aggregations
Produces the mart being delivered
→
Foundation
Zero-Copy Architecture
Databricks (Spark) attaches to the same Iceberg tables
→
Intelligence
Query & Analytics Engine
Read path for the Redshift load job
→
Activation
Databricks / Redshift
Spark/Iceberg attachment (Databricks) or scheduled load job (Redshift)

💼 Business Context

  • Data science teams standardized on Databricks get direct Spark access to the same tables everything else reads, without a duplicate ML-specific copy
  • Legacy Redshift-based BI teams remain supported during a migration window without blocking their existing dashboards
  • Owned by Data Engineering / Data Science Platform

🔌 Technical Overview

Databricks connects to the lakehouse the same way any Iceberg-compatible Spark engine does — direct attachment to Azure Data Lake Storage Gen2 tables, no export job required, matching the pattern already established for ad hoc Spark access under Zero-Copy Architecture. Redshift, without native Iceberg support, receives a scheduled load job (a Docker container on Azure Container Apps Jobs) that reads via the Query & Analytics Engine and writes into Redshift-native tables, used primarily to support legacy dashboards during migration to the lakehouse-native path.

Delivery Modes

Databricks — Spark/Iceberg attachment (zero-copy) Redshift — scheduled load job Used for ML workloads & legacy BI

💾 Databricks Attachment Example

# Databricks notebook, direct Iceberg read — no CXOS export job involved
df = spark.read.format("iceberg") \
    .load("abfss://lakehouse@cxosdata.dfs.core.windows.net/marts/order_fact")

🔗 Integration Points

  • Zero-Copy Architecture — the same principle extended to Databricks' Spark engine
  • Query & Analytics Engine — read path for the Redshift scheduled load
  • Governance & Security — grants and reviews Databricks workspace access to lakehouse tables
  • Predictive Models — Databricks is a common environment for ad hoc feature exploration ahead of formal dbt Python model changes

🧰 Services Consumed

  • Owning microservice — Cxos.Connectors.BatchExport (see the Full Application Service Map)
  • No dedicated database — stateless connector (see Platform Connectors above)

⚠️ Non-Functional Considerations

  • Scale: Databricks attachment has no export-size ceiling; Redshift loads are sized like any scheduled batch export
  • Latency: Databricks reflects the lakehouse near-real-time; Redshift lags by its load schedule
  • Reliability: the Databricks path has no separate copy to go stale; the Redshift path is monitored like any other batch export job
  • Security/Privacy: Databricks workspace access is governed and audited the same as any direct lakehouse consumer, not treated as an external export

🎯 Enterprise Example

A data science team exploring a new feature for the next Predictive Models iteration attaches Databricks directly to curated tables for ad hoc exploration, while the BI team's existing Redshift dashboards keep working unchanged off a nightly load job until their planned migration to querying the lakehouse directly.

← Back to Batch / File Exports