Data Sources → Files & Integrations

CSV / JSON / Parquet

Bulk file-based data — historical exports, partner data drops, and legacy system extracts — landed and ingested on a schedule.

High-Level Design

File drop → Azure Blob Storage → .NET Core batch ingestion on Azure.

Data Source
CSV / JSON / Parquet
Bulk exports from partners or legacy systems
→
Ingestion
.NET Core File Ingestion Service
Azure Blob Storage landing zone → Cxos.Ingestion.Client
→
Processing
.NET Core Batch Worker
Schema inference, validation via Azure Functions
→
Foundation
Data Lakehouse
Landed in raw zone, promoted to curated
→
Intelligence
.NET Core Analytics API
Backfilled into unified profile
→
Activation
.NET Core Activation API
Batch-driven segment refresh

💼 Business Context

  • Not every data source has an API — file drops remain the practical integration path for many legacy and partner systems
  • Enables historical backfill when a new source is onboarded, so CXOS isn't starting from zero
  • Owned by Data Engineering, on behalf of whichever business team owns the source system

🔌 Technical Overview

Files are landed in a dedicated Azure Blob Storage container (one prefix per partner/source). A Blob-triggered .NET Core Azure Function validates the file against a registered schema (CSV header, JSON Schema, or Parquet schema), then streams records through Cxos.Ingestion.Client into the Ingestion API in batches — the same contract as every real-time source. Malformed rows are quarantined to a dead-letter container for manual review.

Supported Formats

CSV JSON / NDJSON Parquet

💾 Sample Ingestion Manifest

{
  "batch_id": "batch_20260801_0300",
  "source": "partner_loyalty_export",
  "format": "parquet",
  "file": "loyalty_2026-08-01.parquet",
  "row_count": 184213,
  "schema_version": "v3",
  "status": "validated"
}

🔗 Integration Points

  • Azure Blob Storage — per-source landing zone
  • Blob-triggered Azure Function (.NET Core) — schema validation + batch ingestion
  • Cxos.Ingestion.Client NuGet package — batch mode
  • Dead-letter Blob container — quarantined malformed records for review

🧰 Services Consumed

  • Owning microservice — Cxos.Connectors.Files (see the Full Application Service Map)
  • No dedicated database — stateless connector (see Platform Connectors above)

⚠️ Non-Functional Considerations

  • Scale: large historical backfills (10M+ rows) are chunked and streamed rather than loaded into memory
  • Latency: file-based ingestion is inherently batch — typically processed within the hour of landing, not real-time
  • Reliability: schema validation happens before ingestion; partial-batch failures don't corrupt already-ingested rows
  • Security/Privacy: landing containers use short retention and Azure Key Vault-managed encryption; PII columns are tagged during schema registration

🎯 Enterprise Example

When onboarding a newly acquired brand, its 3 years of loyalty history arrive as a Parquet export. The Blob-triggered ingestion function validates and streams 2M+ historical records into the Data Lakehouse overnight, letting the Identity API backfill unified profiles before the brand's storefront cuts over to CXOS-powered personalization.

← Back to Files & Integrations