Predictive Models
The shared framework for training, versioning, and serving machine learning models — churn, LTV forecasting, demand — that other AI & Insights capabilities build on.
High-Level Design
Predictive Models is the shared ML platform every scoring capability in this module is built on.
💼 Business Context
- Gives every predictive use case (churn, LTV, propensity) a shared, auditable training and deployment pipeline instead of one-off notebooks nobody can reproduce
- Makes model performance and drift visible to the business, not just to the data science team that built it
- Owned by Data Science, feature pipelines co-owned with Data Engineering
🔌 Technical Overview
Feature engineering happens as dbt Python models — version-controlled and tested like any other dbt model — producing curated feature tables (marts.ml_features) directly from the lakehouse. Azure Machine Learning handles training, experiment tracking, and model registry; trained models are deployed behind a .NET Core Analytics/AI API scoring endpoint (a Docker container on Azure Container Apps) so downstream capabilities like Propensity Scores call a stable API rather than loading a model file directly.
Pipeline Stages
💾 Model Registry Entry
{
"model_name": "churn_risk_v4",
"framework": "azureml",
"trained_on": "marts.ml_features (as of 2026-07-28)",
"auc": 0.87,
"deployed_endpoint": "/v1/score/churn_risk_v4"
}
🔗 Integration Points
- Rollups & Aggregations / dbt Python models — feature engineering pipeline
- Azure Machine Learning — training, experiment tracking, model registry
- Propensity Scores — the primary consumer of trained model outputs
- Application Insights — tracks scoring-endpoint latency and error rates in production
🧰 Services Consumed
- Owning microservice —
Cxos.Intelligence.Api(see the Full Application Service Map) - Database — Azure Data Explorer/Kusto (scoring time series) + Azure Database for PostgreSQL
⚠️ Non-Functional Considerations
- Scale: training runs on a scheduled cadence (not per-request), so cost scales with retraining frequency, not query volume
- Latency: the scoring API returns predictions in under 100ms per customer, using a pre-loaded model rather than cold-starting per request
- Reliability: every deployed model is versioned in the registry, so a regression can be rolled back to the prior version without retraining
- Security/Privacy: feature tables inherit classification tags from their source fields — a model trained on PII-derived features is itself treated as a sensitive artifact
🎯 Enterprise Example
A new churn model scores 0.87 AUC in offline evaluation but drifts to 0.79 in production after a promotional pricing change shifts customer behavior. Drift monitoring flags the gap, and the team rolls back to the prior registry version while retraining — a controlled response instead of silently serving degraded predictions.