top of page

Pilot-to-Production Readiness Checklist for AI Features

Shawn West
7 hours ago
5 min read

The pilot ended with a standing ovation. Six weeks later the feature was switched off, and nothing about the model had changed.


The pilot was a claims-triage assistant at a regional insurer. For eight weeks a small team watched it sort incoming claims into fast-track and review queues, and the numbers were good: adjusters said it saved them time, and the demo to the steering committee went well. Production approval followed within a week.


Then it met Monday. The integration that had read claims from an export file now had to read them from the core claims system, which needed a security review nobody had scheduled. The prompt engineer who had tuned it went back to their own team. Nobody owned the weekly quality review, so when a policy change shifted the mix of claims, accuracy slid for three weeks before an adjuster complained. When finance asked what the feature had saved, the only numbers available were usage counts.


(Composite example, assembled from patterns that recur when AI pilots move into production. The organisation is not a real one.)


None of that is a model failure. Every missing piece was operational work that the pilot was never asked to prove. This checklist puts those pieces in front of the decision, before the scaling budget is committed rather than after.


Why a good pilot is weak evidence


A pilot answers one question well: can this work under supervision, on a curated slice of data, with its builders watching? Production asks different ones. Who notices when it degrades? Who can change it? What does it cost per task at real volume? Who is accountable for the business result?


Industry analysts expect a large share of these projects to stall. Gartner predicted in June 2025 that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls (Gartner, 2025). Notice that none of the three reasons is the model. They are the operating work this checklist scores. The deeper argument is in the pilot-to-production gap is an org problem, not a model problem.


Part 1: the five capabilities production needs


Score each from 0 (nothing in place) to 3 (production-grade). Score what exists today, not what is planned.


#

Capability

What "ready" looks like

Score (0–3)

1

Monitoring and observability

Every request, tool call and output is logged well enough to reconstruct a bad interaction, with alerts on quality and cost


2

Evaluation tooling

An evaluation set built from real cases, run automatically on every change to the prompt, model or retrieval data, and able to block a release


3

Operational staffing

A named person owns incidents, the quality review, prompt changes and user reports, as part of their job rather than on the side


4

Integration with real systems

It reads from and writes to the systems users actually work in, through security-reviewed interfaces, not an export file


5

Domain data pipeline

A funded, ongoing process keeps the retrieval content, examples and evaluation cases current



Total: ___ / 15. Our rule of thumb, not an external benchmark: below 10 means you have a pilot, not a production plan.


Part 2: the five ownership gates


Each must be true before you scale. Any false answer names the likely point of failure.


  • ☐ Budget. The production work (infrastructure, operations, integration, data and compliance) is funded as a multi-quarter programme. If the model licence is the biggest line, the plan is missing something.

  • ☐ Owner. One named person owns this workflow's outcome, with both the authority to change scope, prompts and rollout and the accountability for the business result.

  • ☐ Owner location. That person sits in the business unit whose work changes, not only in IT or engineering.

  • ☐ Scope. Production is the same single workflow the pilot proved. Expansion waits until production works.

  • ☐ Success criteria. There is a specific workflow with specific, measurable success criteria, not a general "AI strategy".


Part 3: the five metrics to commit to before launch


Pilots tend to report engagement. Production needs outcomes. Write a target for each line before launch, so the numbers can't be argued into fitting afterwards.


Metric type

What to measure

Target (set before launch)

Outcome

The workflow result: cycle time, error rate or quality, measured before and after


Quality

Sampled human review and evaluation scores on production inputs, not the pilot set


Reliability

Task completion and failure rates, tracked the way an SRE team tracks service-level objectives


Cost

Full cost per task (model, engineering and operations), watched as volume grows


User experience

Satisfaction, complaint rate and override rate: how often a person has to step in



Stop reporting these as success measures: number of queries, number of users, "AI-assisted tasks", and self-reported time saved. They describe activity, not value.


How to use it


Run it as a 60-minute session with the business owner, the engineering lead and whoever will be on call for the feature. Score Part 1 together and argue about the evidence for each number; the arguments are the useful part. Then read Part 2 aloud and answer each gate yes or no. Finally, fill in the Part 3 targets before anyone leaves the room.


The verdict:


  • Green (scale it): 13 or more on Part 1, all five gates true, all five targets written down.

  • Yellow (fix first): gaps in one or two areas. Fund and close them before scaling.

  • Red (don't scale yet): no owner, no budget plan, or only engagement metrics.


(These thresholds are ShiftQuality's recommendation for structuring the decision. Adjust them to your risk appetite, but set them before you score.)


If the feature is an agent that takes actions rather than one that answers questions, run this alongside the autonomy ladder in what AI agent management actually involves, because a pilot that worked at "draft" says little about "apply". And if the scores are already low, the five ways agents fail in production tells you which gap is likely to bite first.


Where this changes


Two exceptions are worth naming. An internal tool used by a handful of experts who can see every output can launch at Yellow, because the users are the monitoring. A feature that makes or influences decisions about people (credit, hiring, claims) should not launch below Green. In the EU, many of those uses are high-risk systems under the AI Act, with obligations applying from 2 December 2027 under the Digital Omnibus.


Final takeaway


A pilot proves the model can do the task. Production proves the organisation can run it. Score the operating work before you scale, and write the targets down while the decision is still yours to make.


Sources


  • Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

  • Breck, E. et al., The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction, IEEE Big Data, 2017. https://research.google/pubs/the-ml-test-score-a-rubric-for-ml-production-readiness-and-technical-debt-reduction/

  • Google, Site Reliability Engineering, chapter 4, "Service Level Objectives". https://sre.google/sre-book/service-level-objectives/

  • NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

  • Regulation (EU) 2024/1689 (the AI Act), as amended by Regulation (EU) 2026/1744, for the Annex III application date. https://eur-lex.europa.eu/eli/reg/2024/1689/oj

bottom of page