Intelligence & Services → AI & Insights

Natural Language Query

Lets business users ask a question in plain English and get back a governed answer, grounded in the Semantic Layer's approved metrics rather than free-form SQL generation.

High-Level Design

Natural Language Query translates English into a semantic-layer request, never raw SQL against the lakehouse.

Data Source
Data Sources
Every touchpoint and business system
→
Ingestion
Ingestion Layer
SDKs, connectors, protocols
→
Processing
Semantic Layer
The approved metric/dimension vocabulary grounding the model
→
Foundation
marts.* tables
Ultimately what the resolved query reads
→
Intelligence
Natural Language Query
Azure OpenAI Service, constrained to Semantic Layer objects
→
Activation
AI Copilot
The primary interface this capability powers

💼 Business Context

  • Lowers the barrier for business users who don't know SQL or the Semantic Layer's exact metric names to still get a governed, correct answer
  • Reduces the volume of one-off 'can you pull me a number' requests that land on Analytics Engineering
  • Owned by Analytics Engineering / Data Science

🔌 Technical Overview

Azure OpenAI Service is used to translate a natural-language question into a structured request against the Semantic Layer's registered metrics and dimensions — never into raw SQL against lakehouse tables directly — so the model's only job is intent parsing and mapping to an approved vocabulary, not query correctness. The .NET Core Analytics/AI API (a Docker container on Azure Container Apps) validates the resolved metric/dimension combination exists and is entitled to the caller before executing it through the normal Query Optimizer / Caching path.

Guardrails

Grounded in Semantic Layer only No raw SQL generation Entitlement check before execution Ambiguous questions ask for clarification

💾 NL Query Resolution

"What was APAC revenue last week?"
  -> resolved: metric=net_revenue, dimension={region:"APAC"}, time_grain=week, range=last_week
  -> executed via Query Optimizer (same path as any dashboard query)

🔗 Integration Points

  • Azure OpenAI Service — natural-language intent parsing
  • Semantic Layer — the constrained vocabulary the model is grounded in
  • Query Optimizer / Caching — executes the resolved query the normal way, no special-cased path
  • AI Copilot — the conversational interface built on top of this capability

🧰 Services Consumed

  • Owning microservice — Cxos.Intelligence.Api (see the Full Application Service Map)
  • Database — Azure Data Explorer/Kusto (scoring time series) + Azure Database for PostgreSQL

⚠️ Non-Functional Considerations

  • Scale: language parsing cost scales with question volume, not data volume, and is bounded by the fixed size of the semantic vocabulary
  • Latency: intent resolution adds roughly 1-2 seconds ahead of normal query execution latency
  • Reliability: because the model only ever selects from pre-registered metrics, a hallucinated or malformed query cannot reach the lakehouse — it fails at the resolution step, not silently returning wrong data
  • Security/Privacy: the resolved query still passes through Access Control per the asking user's entitlements — natural language is not a privilege-escalation path

🎯 Enterprise Example

A regional manager types "how did APAC do last week compared to the week before" instead of building a dashboard filter. The system resolves it to two semantic-layer queries against net_revenue and returns a governed, cache-eligible answer in seconds — the same number a formal dashboard would show.

← Back to AI & Insights