Data Governance
What is the AI allowed to know?
Everything an AI system can read, retrieve, memorize, or learn from: access boundaries, classification and de-identification, residency and retention, lineage from source to model to output, and the provenance of training and grounding data. GenAI turned every document store into potential model input — so data governance became the substrate of AI governance.
AI turns latent data-governance debt into live exposure at conversational speed.
Why leadership should care
- Retrieval-augmented AI surfaces whatever permissions allow — including the over-shared folders nobody audited for a decade. The model is rarely the leak; the permissions are.
- Training-data provenance is now a balance-sheet issue: the Bartz v. Anthropic settlement priced pirated training books at $1.5B.
- Residency and retention promises made to regulators must survive contact with caching, logging, and vendor abuse-monitoring — details most AI contracts never mention.
- Which data classes may reach which AI systems — and which may never (crown jewels, regulated categories)?
- What are our residency and retention red lines, contractually and technically?
- Do we require no-training terms and IP indemnity from every AI vendor?
Sensitive Data Protection (Cloud DLP)
GASecurity
Discovery, classification, and de-identification across 150+ infotypes.
PII/PHI kept out of training data, prompts, responses, and logs.
VPC Service Controls
GASecurity
Service perimeters around AI APIs block exfiltration paths.
Keeps AI workloads inside a data perimeter, including partner-model calls.
Customer-managed encryption keys
GASecurity
Customer key custody over platform resources, tuning artifacts, agent state.
Key control and crypto-shredding for regulated and sovereign workloads.
Data residency & zero-data-retention
GAGemini Enterprise Agent Platform
Regional endpoints for in-region processing; ZDR by disabling the 24h cache.
Residency and retention posture as explicit configuration, per model.
AI/ML Privacy Commitment
GAContractual
Customer data is not used to train Google models without permission.
The first procurement blocker, answered contractually rather than by policy blog.
Knowledge Catalog (formerly Dataplex Universal Catalog)
GAData governance
Unified catalog and automatic lineage across data and AI assets.
The map auditors ask for: which data trained which model behind which decision.
Access Transparency & Access Approval
GASecurity
Logs — and approval gates — for Google-personnel access to customer content.
Provider-insider assurance regulators increasingly ask about.
Google Distributed Cloud air-gapped
GASovereign
Gemini on-prem with no connectivity to Google; IL5/IL6-class isolation.
Sovereign and classified AI where data can never leave the perimeter.
Cards link to official documentation. Status is a snapshot (August 2026) — verify per component before contractual commitments. Full mapping and honest gaps: 08 · Google Cloud.
An AI system inherits every data-governance debt you already have — then exposes it through a chat box.
- The Copilot-era lesson: enterprise search exposes permission debt. Fix sharing before you index, or the AI will audit your ACLs for you — in production.
- Wells Fargo ran a quarter-billion assistant interactions with zero PII reaching the LLM — tokenize first is an architecture, and it's copyable.
- Ask where inference runs, not just where data is stored. At-rest residency with US inference is a common surprise.
- Zero data retention is a configuration, not a default — caching and abuse-monitoring settings decide your real posture.
- What data could your AI assistants reach today that you would not show a new hire on day one?
- When did you last review sharing permissions on the corpora your AI retrieves from?
- Can PII reach your model provider — in prompts, logs, or caches — and would you know?
- For your most important model: what data trained or tuned it, and could you prove rights to it?
- What did you contractually agree about training, retention, and residency with each AI vendor?
Wells Fargo
Google CloudFinancial Services
You don't need to trust the model with PII to use the model — that's an architecture decision.
Highmark Health
Google CloudHealthcare
Centralize access and measure usage first — governance data is what lets you expand safely.
US Air Force
MarketGovernment / Defense
Don't ban shadow AI; out-compete it, measure it, then institutionalize it.