CSV / JSON / Parquet
Bulk file-based data — historical exports, partner data drops, and legacy system extracts — landed and ingested on a schedule.
High-Level Design
File drop → Azure Blob Storage → .NET Core batch ingestion on Azure.
💼 Business Context
- Not every data source has an API — file drops remain the practical integration path for many legacy and partner systems
- Enables historical backfill when a new source is onboarded, so CXOS isn't starting from zero
- Owned by Data Engineering, on behalf of whichever business team owns the source system
🔌 Technical Overview
Files are landed in a dedicated Azure Blob Storage container (one prefix per partner/source). A Blob-triggered .NET Core Azure Function validates the file against a registered schema (CSV header, JSON Schema, or Parquet schema), then streams records through Cxos.Ingestion.Client into the Ingestion API in batches — the same contract as every real-time source. Malformed rows are quarantined to a dead-letter container for manual review.
Supported Formats
💾 Sample Ingestion Manifest
{
"batch_id": "batch_20260801_0300",
"source": "partner_loyalty_export",
"format": "parquet",
"file": "loyalty_2026-08-01.parquet",
"row_count": 184213,
"schema_version": "v3",
"status": "validated"
}
🔗 Integration Points
- Azure Blob Storage — per-source landing zone
- Blob-triggered Azure Function (.NET Core) — schema validation + batch ingestion
- Cxos.Ingestion.Client NuGet package — batch mode
- Dead-letter Blob container — quarantined malformed records for review
🧰 Services Consumed
- Owning microservice —
Cxos.Connectors.Files(see the Full Application Service Map) - No dedicated database — stateless connector (see Platform Connectors above)
⚠️ Non-Functional Considerations
- Scale: large historical backfills (10M+ rows) are chunked and streamed rather than loaded into memory
- Latency: file-based ingestion is inherently batch — typically processed within the hour of landing, not real-time
- Reliability: schema validation happens before ingestion; partial-batch failures don't corrupt already-ingested rows
- Security/Privacy: landing containers use short retention and Azure Key Vault-managed encryption; PII columns are tagged during schema registration
🎯 Enterprise Example
When onboarding a newly acquired brand, its 3 years of loyalty history arrive as a Parquet export. The Blob-triggered ingestion function validates and streams 2M+ historical records into the Data Lakehouse overnight, letting the Identity API backfill unified profiles before the brand's storefront cuts over to CXOS-powered personalization.