Intelligence & Services → AI & Insights

Predictive Models

The shared framework for training, versioning, and serving machine learning models — churn, LTV forecasting, demand — that other AI & Insights capabilities build on.

High-Level Design

Predictive Models is the shared ML platform every scoring capability in this module is built on.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
dbt Python Models
Feature prep (Rollups & Aggregations)
→
Foundation
marts.ml_features
Curated feature tables in the lakehouse
→
Intelligence
Predictive Models
Azure Machine Learning training + registry
→
Activation
Propensity Scores, AI Copilot
Consumers of trained model outputs

💼 Business Context

  • Gives every predictive use case (churn, LTV, propensity) a shared, auditable training and deployment pipeline instead of one-off notebooks nobody can reproduce
  • Makes model performance and drift visible to the business, not just to the data science team that built it
  • Owned by Data Science, feature pipelines co-owned with Data Engineering

🔌 Technical Overview

Feature engineering happens as dbt Python models — version-controlled and tested like any other dbt model — producing curated feature tables (marts.ml_features) directly from the lakehouse. Azure Machine Learning handles training, experiment tracking, and model registry; trained models are deployed behind a .NET Core Analytics/AI API scoring endpoint (a Docker container on Azure Container Apps) so downstream capabilities like Propensity Scores call a stable API rather than loading a model file directly.

Pipeline Stages

dbt Python feature models Azure ML training & registry Scoring API (Docker/ACA) Drift monitoring

💾 Model Registry Entry

{
  "model_name": "churn_risk_v4",
  "framework": "azureml",
  "trained_on": "marts.ml_features (as of 2026-07-28)",
  "auc": 0.87,
  "deployed_endpoint": "/v1/score/churn_risk_v4"
}

🔗 Integration Points

  • Rollups & Aggregations / dbt Python models — feature engineering pipeline
  • Azure Machine Learning — training, experiment tracking, model registry
  • Propensity Scores — the primary consumer of trained model outputs
  • Application Insights — tracks scoring-endpoint latency and error rates in production

🧰 Services Consumed

  • Owning microservice — Cxos.Intelligence.Api (see the Full Application Service Map)
  • Database — Azure Data Explorer/Kusto (scoring time series) + Azure Database for PostgreSQL

⚠️ Non-Functional Considerations

  • Scale: training runs on a scheduled cadence (not per-request), so cost scales with retraining frequency, not query volume
  • Latency: the scoring API returns predictions in under 100ms per customer, using a pre-loaded model rather than cold-starting per request
  • Reliability: every deployed model is versioned in the registry, so a regression can be rolled back to the prior version without retraining
  • Security/Privacy: feature tables inherit classification tags from their source fields — a model trained on PII-derived features is itself treated as a sensitive artifact

🎯 Enterprise Example

A new churn model scores 0.87 AUC in offline evaluation but drifts to 0.79 in production after a promotional pricing change shifts customer behavior. Drift monitoring flags the gap, and the team rolls back to the prior registry version while retraining — a controlled response instead of silently serving degraded predictions.

← Back to AI & Insights