Destinations & Activation → Batch / File Exports

S3 / GCS / Azure Blob

Scheduled bulk export of lakehouse data to a customer or partner's own cloud object storage — the simplest, most universal batch destination.

High-Level Design

The lowest-friction batch destination: files land in a bucket the receiving team already owns.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Rollups & Aggregations
Produces the mart the export is sourced from
→
Foundation
marts.* tables
Analytics-ready export source
→
Intelligence
Query & Analytics Engine
Executes the export query
→
Activation
S3 / GCS / Azure Blob
Scheduled .NET Core export job (Azure Functions Timer)

💼 Business Context

  • The lowest-friction destination for partners and internal teams who already have their own cloud storage and just need the data delivered there on a schedule
  • Avoids building a bespoke integration per partner when a file drop is sufficient
  • Owned by Data Engineering / Partner Integrations

🔌 Technical Overview

A scheduled Azure Functions Timer job (a .NET Core service packaged as a Docker container) queries the relevant marts table via the Query & Analytics Engine, writes the result as Parquet or CSV, and uploads it to the destination bucket — Amazon S3, Google Cloud Storage, or Azure Blob Storage — using the appropriate cloud SDK and a partner-scoped credential stored in Azure Key Vault. Each export run is logged with row count and checksum for downstream validation.

Destinations

Amazon S3 Google Cloud Storage Azure Blob Storage Parquet / CSV format

💾 Export Job Manifest

{
  "destination": "s3://partner-northwind/cxos-exports/",
  "table": "marts.order_fact",
  "format": "parquet",
  "schedule": "0 2 * * *",
  "row_count": 481200,
  "checksum": "sha256:9f2a..."
}

🔗 Integration Points

  • Query & Analytics Engine — executes the export query against marts tables
  • Azure Key Vault — stores partner-scoped destination credentials
  • Azure Functions (Timer trigger) — runs the scheduled export job
  • Audit Logs — records every export run for compliance traceability

🧰 Services Consumed

  • Owning microservice — Cxos.Connectors.BatchExport (see the Full Application Service Map)
  • No dedicated database — stateless connector (see Platform Connectors above)

⚠️ Non-Functional Considerations

  • Scale: export size scales with the requested table/date range, not total lakehouse size, since each export is scoped by query
  • Latency: batch by design — runs on a defined schedule (typically daily/hourly), not on-demand
  • Reliability: checksum and row-count validation lets the receiving side confirm a complete, uncorrupted transfer before consuming it
  • Security/Privacy: only fields the export configuration explicitly includes are written — the export job itself is entitlement-scoped like any other consumer

🎯 Enterprise Example

A logistics partner needs daily order data to plan fulfillment capacity. A nightly export job drops a Parquet file into their S3 bucket by 2am local time, replacing what used to be a manual weekly CSV emailed by an analyst.

← Back to Batch / File Exports