Data Classification
Tags every field with a sensitivity level (public, internal, PII, sensitive-PII) — the foundation every other privacy and security control builds on.
High-Level Design
Nothing else in Governance & Security works without classification first.
💼 Business Context
- Nothing else in Governance & Security works without this — access control, encryption scope, and consent enforcement all key off classification
- Gives the business a defensible, consistent answer to "how do you know where your sensitive data is"
- Owned by Data Governance, enforced at Schema Registry registration time
🔌 Technical Overview
Every field is classified at schema registration time (enforced by the Event Schema & Registry's Governance Rules) using a standard taxonomy: public, internal, PII, and sensitive-PII (financial, health). Classification tags are stored in the Catalog alongside each field's technical schema and propagate automatically to derived fields computed from a classified source — a field derived from an email address inherits the PII tag rather than starting unclassified.
Classification Levels
💾 Classification Tag
{
"field": "curated.customer_profile.email",
"classification": "pii",
"subcategory": "contact_info",
"classified_by": "schema-registry-governance-rules",
"classified_at": "2026-01-15T00:00:00Z"
}
🔗 Integration Points
- Event Schema & Registry's Governance Rules — enforces classification at registration time
- Catalog — stores classification alongside technical schema
- Access Control (RBAC/ABAC) — masking rules are driven by classification
- Encryption — sensitive-PII may warrant stronger encryption keys/rotation policy
🧰 Services Consumed
- Owning microservice —
Cxos.Foundation.Api(see the Full Application Service Map) - Database — ADLS Gen2 (Iceberg) + Azure Database for PostgreSQL (policy/retention state)
⚠️ Non-Functional Considerations
- Scale: classification is metadata, not a runtime computation — no performance concern
- Latency: not applicable — classification is a design-time and registration-time control
- Reliability: automatic propagation to derived fields prevents classification gaps as new fields are computed
- Security/Privacy: the foundational control — every other privacy/security mechanism in this handbook depends on classification being correct
🎯 Enterprise Example
A new derived field, customer_email_domain, is computed from the classified email field. Because classification propagates automatically, it correctly inherits at least an 'internal' sensitivity tag rather than defaulting to unclassified and silently bypassing access controls built around classification.