top of page

AI Governance in the Pipeline: A Checklist

Shawn West
1 day ago
4 min read

The model was approved in March. The model that was running in September had never been reviewed by anyone.


A lending team's credit-risk model went through a careful governance review: a responsible-AI board read the documentation, asked good questions about subgroup performance and approved it with conditions. Over the following six months the model was retrained four times as new data arrived. Each retrain went out through the normal deployment pipeline, and the pipeline checked that the service started and responded. Nobody checked whether the approval conditions still held. When an analyst finally compared approval rates by applicant group, the gap the board had asked to keep under watch had widened after the second retrain.


(Composite example, assembled from patterns that recur in ML governance. The organisation is not a real one.)


The board did its job. The trouble is that it reviewed a point in time, and the model changed through a channel the board couldn't see. Governance that only exists in a meeting can't follow a system that changes on every merge. This checklist moves the controls into the delivery pipeline, where every change has to pass them.


Theatre versus governance that works


Governance theatre

Governance in the pipeline

A quarterly board reviews a slide deck

A CI gate blocks a deploy automatically

A model card is written as a PDF at final review

The model card is version-controlled next to the code

One-time evaluation before launch

Evaluation on every change, plus production monitoring

Evidence lives in someone's memory

Evidence lives in the commit history and pipeline logs


Governance that works has three properties: it is automated where it can be, embedded in the way the team already ships, and it produces evidence that each decision was deliberate.


Pillar 1: documentation as code


  • ☐ The model card lives alongside the model code and is version-controlled. The format comes from Mitchell et al.'s Model Cards for Model Reporting (2019).

  • ☐ The pipeline cannot deploy without a model card that meets a minimum standard.

  • ☐ The card covers training data, evaluation results broken down by relevant subgroups, known limitations, intended use and out-of-scope use.

  • ☐ The card changes in the same pull request as the model, so it describes what is running.


Test you can run: open the model card for the version in production now. If its last change is older than the model's last change, the card is describing a different model.


Pillar 2: evaluation gates that block a release


Any one of these failing should stop the deploy:


  • ☐ Accuracy at or above an agreed threshold on a held-out evaluation set.

  • ☐ Fairness metrics within agreed bounds across predefined subgroups, for example the false-positive rate per group.

  • ☐ Prediction distribution consistent with the previous version. (Illustration: a model that approved about 60% of applications last release and approves about 30% now should stop and be explained, even if accuracy looks fine.)

  • ☐ Data validation: row counts, feature distributions and null rates within expected ranges.


Google's ML Test Score rubric (Breck et al., 2017) is a useful source of further checks. Gates won't catch everything. They catch the obvious problems every time, so people can spend their attention on the subtle ones. Quality gates for shipping AI covers how to set the thresholds.


Pillar 3: production monitoring with governance hooks


  • ☐ Prediction drift: alert when the output distribution moves away from the baseline.

  • ☐ Fairness drift: track subgroup performance over time, not just at launch.

  • ☐ Feature drift: alert when production inputs diverge from training inputs.

  • ☐ Feedback: route complaints, override rates and downstream outcomes back to the model team. (Illustration: if reviewers override two in five recommendations, the model isn't doing the job it was approved for.)

  • ☐ Automated response: an anomaly alerts the model team, escalates to the governance owner, and can roll back to the last known-good version.


Who owns what


Owner

Owns

Model team

The model card, evaluation results and training-data documentation

Platform team

The evaluation gates, the deployment pipeline and the monitoring infrastructure

Governance function (board, risk team or responsible-AI lead)

The policies, thresholds and escalation routes: what the gates check


This split avoids the two usual failures: governance as someone else's problem, and governance as a bottleneck. Automate the routine checks, and keep human review for new use cases, new data, high-risk deployments and policy questions. For AI agents, which change through configuration as well as code, what AI agent management involves extends this to autonomy levels and permissions.


Where the regulations land


The EU AI Act asks providers of high-risk systems for technical documentation, record-keeping, human oversight and post-market monitoring. Those obligations apply to the stand-alone Annex III systems from 2 December 2027, after the Digital Omnibus moved the date (Regulation (EU) 2026/1744). The transparency duties in Article 50 already apply from 2 August 2026. What the EU AI Act deadline actually means has the full timeline.


A pipeline built this way produces most of that evidence as a by-product: version-controlled model cards, timestamped evaluation results, gate audit trails and monitoring records. The NIST AI Risk Management Framework's MEASURE and MANAGE functions describe the same activities in non-regulatory terms.


How to use it


Take one model or AI feature that is in production today. Walk the three pillars with its model team and platform lead, and mark each item as in place, partial or missing. Then answer one question for the governance owner: if a regulator or an affected customer asked how this version got deployed, could we answer from records alone? Each missing item is a line on next quarter's plan.


Final takeaway


A review board governs the version it saw. A pipeline governs every version that ships. Put the checks where changes happen, and governance stops depending on whether anyone remembered to call a meeting.


Sources


  • Mitchell, M. et al., Model Cards for Model Reporting, FAT* 2019. https://arxiv.org/abs/1810.03993

  • Breck, E. et al., The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, IEEE Big Data, 2017. https://research.google/pubs/the-ml-test-score-a-rubric-for-ml-production-readiness-and-technical-debt-reduction/

  • NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

  • Regulation (EU) 2024/1689 (the AI Act), Articles 11, 12, 14, 50 and 72, as amended by Regulation (EU) 2026/1744. https://eur-lex.europa.eu/eli/reg/2024/1689/oj

bottom of page