Workflow Engine
Orchestrates multi-step operational processes — approval flows, escalations, scheduled jobs — across Intelligence & Services and beyond, distinct from real-time customer journeys.
High-Level Design
Workflow Engine coordinates internal operational steps, not customer-facing journeys.
💼 Business Context
- Gives operational processes — a schema-governance approval, an anomaly escalation, a scheduled retention job — a consistent, auditable state machine instead of ad hoc scripts and tribal knowledge
- Makes it possible to answer "where is this approval stuck" without asking the one engineer who remembers
- Owned by Platform Engineering
🔌 Technical Overview
The Workflow Engine is a .NET Core state-machine service (Docker container on Azure Container Apps) that models each operational process as a sequence of steps with defined transitions, using Azure Service Bus topics to advance state on completion of an async step (a job finishing, an approver responding) rather than polling. It is the same orchestration layer scheduled Governance & Security jobs (Lifecycle Management, deletion requests) and Data Quality Monitoring escalations run under, giving every operational process the same visibility model.
Workflow Types
💾 Workflow State
{
"workflow_id": "wf_88213",
"type": "schema_change_approval",
"state": "pending_approver",
"steps_completed": ["submitted", "governance_review"],
"next_step": "data_owner_approval"
}
🔗 Integration Points
- Azure Service Bus — topic/queue backbone advancing workflow state
- Lifecycle Management, Event Schema & Registry — operational jobs orchestrated through this engine
- Alerts & Notifications — subscribes to workflow state-change events
- Application Insights — traces each workflow's step-by-step execution
🧰 Services Consumed
- Owning microservice —
Cxos.Operations.Api(see the Full Application Service Map) - Database — Azure Database for PostgreSQL + Azure Cosmos DB Table API + Azure Data Explorer
⚠️ Non-Functional Considerations
- Scale: workflow volume is operational (dozens to hundreds concurrently), several orders of magnitude below customer event volume
- Latency: state transitions occur within seconds of the triggering event via Service Bus, not on a polling delay
- Reliability: workflow state is persisted at every transition, so a service restart resumes in-flight workflows rather than losing them
- Security/Privacy: workflow definitions and history are internal-operational data, access-controlled by role rather than customer-data classification
🎯 Enterprise Example
A proposed breaking schema change is submitted for approval. The Workflow Engine routes it through governance review and data-owner sign-off, and anyone can check its exact state at any time — replacing what used to be a Slack thread nobody could reliably locate two weeks later.