Skip to content
Layer 3The governance stack/ Models

Model Governance

Which models do we trust — and how do we know?

Choosing, validating, and tracking the models themselves: a curated catalog with an approval path, provenance checks on third-party and open weights, evaluation as a repeatable gate rather than a launch ritual, version pinning against silent provider churn, and bias and safety testing proportionate to what the model decides. Classic model risk management supplies the skeleton; GenAI forces it to evaluate behavior distributions, not fixed test vectors.

The risk, in one line

An unevaluated model in production is an unread contract you signed on behalf of the business.

Why leadership should care

  • Frontier models are third-party: weights, training data, and alignment process are not inspectable. Trust must come from evaluation, contract, and provenance — not inspection.
  • Providers update hosted models continuously; behavior shifts with no change ticket on your side. Version pinning and regression evals are the only counterweight.
  • Where models touch decisions about people, bias liability is live now: courts allowed a nationwide collective action over AI hiring screens and held that vendors can be liable as employers' agents.
1,598
court decisions involving AI-fabricated citations tracked by mid-2026 — unevaluated output has legal consequences
Decisions only leadership can make
  • Do we run one vetted model catalog with an approval path — or do teams pick models ad hoc?
  • What evaluation bar must any model clear before production, and who owns the bar?
  • What is our position on open-weight models, and who governs them once downloaded?
Read the claims right
RegulationEU AI Act GPAI transparency (Art. 53)PracticeNIST GenAI Profile (AI 600-1)PracticeSR 11-7 model risk lineageStandardISO/IEC 42001 lifecycle controls
In the room — discussion points

You no longer validate an artifact you built — you continuously evaluate behavior you rent.

  • Morgan Stanley's famous move wasn't the chatbot — it was writing the evals before the rollout. Evaluation discipline is the control regulators will ask to see.
  • One vetted catalog beats per-team model choice: Goldman and Walmart both built exactly this before scaling.
  • Silent model churn is real: providers update behavior under your feet. Pin versions and make updates rerun your regression suite.
  • Open models move governance onto you: once weights are downloaded, platform guardrails and logging no longer apply.
Questions to ask your organization
  • How many distinct models are in production, and who approved each one?
  • What evaluation evidence exists for your most business-critical AI use case?
  • What happens on your side when a provider updates a hosted model?
  • Do any models influence decisions about people — and when were they last bias-tested?
  • Who governs open-weight models your teams have downloaded?