Transformation & Processing → Batch Processing

Machine Learning Feature Prep

Builds and maintains the feature tables ML models train on and score against — the bridge between raw data and the AI & Insights layer.

High-Level Design

Feature Prep is where most of the real work behind "AI-powered" personalization happens.

Data Source
Data Lakehouse (curated zone)
Source event data
→
Ingestion
dbt Python Models
pandas/scikit-learn feature transforms
→
Processing
ML Feature Prep
dbt models (SQL + Python) — feature engineering
→
Foundation
Feature Store Table
Data Lakehouse — computed features
→
Intelligence
AI & Insights API
Reads features at inference time
→
Activation
Azure Machine Learning
Training jobs read point-in-time snapshots

💼 Business Context

  • Model quality is bounded by feature quality — this is where most of the real work in 'AI-powered' personalization actually happens
  • Reusable feature tables mean a new model doesn't start from raw events every time — it builds on a shared, governed feature set
  • Owned by Data Science / ML Engineering, using Data Engineering's batch infrastructure

🔌 Technical Overview

Most feature tables are ordinary SQL dbt models (recency/frequency/monetary aggregates); features that need array or ML-library logic (e.g., a category-affinity vector) use dbt's Python model support, running pandas/scikit-learn inside the same orchestrated pipeline instead of a separate bespoke feature-pipeline codebase. Every feature model gets the same version control, dbt test data-quality checks, and automatic lineage as any other transformation. The .NET Core scheduler, packaged as a Docker container on Azure Container Apps Jobs, triggers the feature-tagged model selection on a schedule; the AI & Insights API reads the latest feature vector at inference time, and Azure Machine Learning training jobs read historical feature snapshots for point-in-time-correct training data.

Example Features

recency_days purchase_frequency_90d category_affinity_vector engagement_score

💾 dbt Python Model — Feature

# models/features/customer_engagement_score.py
def model(dbt, session):
    df = dbt.ref("stg_customer_events").to_pandas()
    df["engagement_score"] = (
        df["session_count_30d"] * 0.4 + df["purchase_count_30d"] * 0.6
    )
    return df[["customer_key", "engagement_score"]]

🔗 Integration Points

  • dbt Core (SQL models) + dbt Python models — feature engineering, version-controlled and tested
  • Feature store table (Data Lakehouse) — where computed features are written
  • AI & Insights API — reads features at inference time
  • Azure Machine Learning — training jobs read point-in-time feature snapshots

🧰 Services Consumed

  • Owning microservice — Cxos.Processing.Api (see the Full Application Service Map)
  • Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)

⚠️ Non-Functional Considerations

  • Scale: dbt's incremental materialization updates only changed customers' feature rows rather than a full recompute each run
  • Latency: features typically refresh daily; models needing fresher signals combine batch features with the real-time enrichment layer
  • Reliability: point-in-time correctness is enforced so training data never leaks future information into a feature computed 'as of' an earlier date
  • Security/Privacy: feature tables are subject to the same governance classification as the raw data they are derived from

🎯 Enterprise Example

The propensity-to-convert model used by AI & Insights trains on a point-in-time-correct dbt feature snapshot rather than today's live data — the same dbt test suite that validates the feature model in production also catches a future-leaking join before it ever reaches a training run.

← Back to Batch Processing