Data Modeling
Transforms raw, event-shaped data into the dimensional/entity models that analytics tools and business users actually think in — customers, orders, products.
High-Level Design
Data Modeling bridges raw events and business-friendly entities.
💼 Business Context
- Business users and BI tools don't think in raw events — they think in customers, orders, and products; modeling bridges that gap
- Establishes a single, governed definition of core business entities instead of every analyst building their own
- Owned by Analytics Engineering, in partnership with each business domain
🔌 Technical Overview
Dimensional models are dbt models following the staging → intermediate → marts layering convention: staging models do light cleanup 1:1 with a source table, intermediate models join and reshape, and marts models are the final fact/dimension tables the Query & Analytics Engine's semantic layer maps business-friendly names onto. dbt's ref() and source() functions build the model dependency graph automatically — this graph is what the Metadata Layer's Lineage capability surfaces, so lineage is a byproduct of writing the models correctly, not separate documentation work. Models materialize as Iceberg tables in the lakehouse by default, or in Snowflake for teams standardized on that warehouse, via dbt's adapter.
Model Types
💾 dbt Model — Dimension
-- models/marts/dim_customer.sql
{{ config(materialized='table') }}
select
customer_key,
tier,
signup_date,
current_timestamp() as valid_from
from {{ ref('stg_customer_profile') }}
🔗 Integration Points
- dbt Core — staging/intermediate/marts model layers, version-controlled alongside application code
- dbt docs — auto-generates the dependency graph that feeds the Metadata Layer's Lineage
- Data Lakehouse analytics-marts zone, or Snowflake via dbt's adapter — model materialization target
- Query & Analytics Engine's semantic layer — maps business terms onto marts models
🧰 Services Consumed
- Owning microservice —
Cxos.Processing.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Cosmos DB Table API (stream checkpoints)
⚠️ Non-Functional Considerations
- Scale: dbt chooses incremental or full-refresh materialization per model based on data volume and change rate, keeping build times manageable as history grows
- Latency: models typically refresh on an hourly-to-daily cadence, matching the rollups' cadence
- Reliability: model changes go through the same pull-request review as an API contract change, since BI tools depend on stability
- Security/Privacy: PII fields in dimension tables are classified and access-controlled the same as any other sensitive data in the lakehouse
🎯 Enterprise Example
When Finance asks 'what was gold-tier customer revenue last quarter', the answer comes from a join between the dim_customer and fct_orders dbt models — reviewed via pull request like any other code change — not a bespoke, unreviewed query against raw events.