Machine Learning Feature Prep
Builds and maintains the feature tables ML models train on and score against — the bridge between raw data and the AI & Insights layer.
High-Level Design
Feature Prep is where most of the real work behind "AI-powered" personalization happens.
💼 Business Context
- Model quality is bounded by feature quality — this is where most of the real work in 'AI-powered' personalization actually happens
- Reusable feature tables mean a new model doesn't start from raw events every time — it builds on a shared, governed feature set
- Owned by Data Science / ML Engineering, using Data Engineering's batch infrastructure
🔌 Technical Overview
Most feature tables are ordinary SQL dbt models (recency/frequency/monetary aggregates); features that need array or ML-library logic (e.g., a category-affinity vector) use dbt's Python model support, running pandas/scikit-learn inside the same orchestrated pipeline instead of a separate bespoke feature-pipeline codebase. Every feature model gets the same version control, dbt test data-quality checks, and automatic lineage as any other transformation. The .NET Core scheduler, packaged as a Docker container on Azure Container Apps Jobs, triggers the feature-tagged model selection on a schedule; the AI & Insights API reads the latest feature vector at inference time, and Azure Machine Learning training jobs read historical feature snapshots for point-in-time-correct training data.
Example Features
💾 dbt Python Model — Feature
# models/features/customer_engagement_score.py
def model(dbt, session):
df = dbt.ref("stg_customer_events").to_pandas()
df["engagement_score"] = (
df["session_count_30d"] * 0.4 + df["purchase_count_30d"] * 0.6
)
return df[["customer_key", "engagement_score"]]
🔗 Integration Points
- dbt Core (SQL models) + dbt Python models — feature engineering, version-controlled and tested
- Feature store table (Data Lakehouse) — where computed features are written
- AI & Insights API — reads features at inference time
- Azure Machine Learning — training jobs read point-in-time feature snapshots
🧰 Services Consumed
- Owning microservice —
Cxos.Processing.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)
⚠️ Non-Functional Considerations
- Scale: dbt's incremental materialization updates only changed customers' feature rows rather than a full recompute each run
- Latency: features typically refresh daily; models needing fresher signals combine batch features with the real-time enrichment layer
- Reliability: point-in-time correctness is enforced so training data never leaks future information into a feature computed 'as of' an earlier date
- Security/Privacy: feature tables are subject to the same governance classification as the raw data they are derived from
🎯 Enterprise Example
The propensity-to-convert model used by AI & Insights trains on a point-in-time-correct dbt feature snapshot rather than today's live data — the same dbt test suite that validates the feature model in production also catches a future-leaking join before it ever reaches a training run.