AI Agents Fail in Production in the Same Five Ways
Updated: 7 hours ago
Rewritten, 4 October 2026. An earlier version of this post led with a figure we could not trace to a reliable source. This version keeps what the evidence supports.
The expense agent passed every demo. Finance had watched it read a receipt, find the matching card transaction, file the claim and post a confirmation in Slack, start to finish, a dozen times. It went live on a Monday.
By Thursday, a controller noticed that some claims marked "filed" didn't exist in the expense system at all. The agent had called the expense API, received a timeout on roughly one call in fifty, and then reported success anyway: it composed the confirmation it would have sent if the call had worked. Nobody's dashboard showed an error, because there wasn't one. There was a completed task and a missing record.
(Composite scenario, drawn from failure patterns that recur in agent deployments. Not a single named incident.)
Nothing in that story is a model failure in the sense people usually mean. The model read receipts accurately. What failed was the system around it, and it failed in a way that is common enough to be predictable. Gartner expects that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls (Gartner, June 2025). That's a forecast, not a measurement. But the reasons it gives are the same ones that show up when you take a failed agent apart.
This piece is for the QA or delivery lead who has to say whether an agent is ready for production. It walks through the five ways agents break, the mechanism behind each, and a gate you can run before launch.
Why the obvious diagnosis is "the model"
When an agent misbehaves, the model is the most visible component and the easiest thing to change. Swapping in a newer model is one line of configuration, it feels like progress, and sometimes it helps. A stronger model really is better at following tool schemas and recovering from unexpected responses.
The trouble is that a model swap leaves the structure untouched. The expense agent would have reported false success on any model, because nothing in the system forced it to check the expense API's answer before claiming success. The five patterns below are properties of the system, which is why they survive model upgrades.
1. Reliability compounds across steps
An agent is a chain of steps: plan, call a tool, read the result, decide, call another. Each step can go wrong, and the chances multiply.
Illustrative calculation: if each step succeeds 95% of the time and a task takes ten steps, the task succeeds about 60% of the time (0.95¹⁰ ≈ 0.60). At 99% per step, ten steps give about 90%, which still means one task in ten ends in an error, a wrong result or a silent failure. Steps aren't perfectly independent in practice, so treat this as a way of seeing the shape, not a forecast.
The demo hides this because it runs a short, familiar path a handful of times. Production runs long, unfamiliar paths thousands of times a day.
Run this: pick your agent's most common task and count its steps from the logs, not from the design document. Estimate each step's success rate honestly from a sample of real runs, then multiply. If the result is below what the business can tolerate, the levers are fewer steps (narrower scope) or deterministic checkpoints between steps. A better prompt won't move a product of ten numbers very far.
2. Tool errors turn into silent success
This is the expense agent's failure, and it's the most damaging pattern because it doesn't look like a failure.
The mechanism: a tool call fails in a way the agent's wrapper doesn't surface clearly. It might be a timeout, an empty response or an error message phrased like content. The model, asked to complete the task, produces the output that success would have produced. Downstream systems take that output at face value.
The fix isn't a cleverer model. It's a rule: the agent's claim is a hypothesis, and the system of record is the truth. If the agent says it filed a claim, the runtime queries the expense system for the claim. If it says it sent an email, the runtime checks the mail log. And a failed tool call should end the task and escalate it, rather than inviting the model to "do its best".
Run this: take a random sample of 50 tasks your agent marked complete last week, and check each one against the system it claims to have changed. Count the mismatches. Any mismatch at all means you're shipping silent failures, and your success-rate dashboard is measuring the agent's opinion of itself.
3. The same input takes a different path
Traditional tests assume that the same input produces the same behaviour. Agents break that assumption by design.
There's published evidence of how large the effect can be, with one important qualification about scope. Moshkovich and colleagues ran a calculator agent on 50 examples, five times each, and measured a mean coefficient of variation of 63% in the structure of its execution paths. They also found 19% variation in accuracy for identical inputs (arXiv 2503.06745, 2025). That's one agent and one task family, not agents in general. It does show that run-to-run variation can be large even for a simple, well-defined task.
What that breaks is any test that asserts how the agent got there: which tool it called first, which branch it took. Those tests flake, get muted, and stop protecting anything.
What to test instead:
Outcomes, not paths. Assert the result and the side effects: the claim exists with the right amount and the right card. Don't assert the route.
Statistical thresholds. Run the same task several times and require a pass rate, such as 19 of 20. Agree the threshold with the business before you see the results.
An adversarial suite. Keep a set of tool timeouts, malformed inputs, expired credentials and hostile instructions. Run it on every prompt, tool or model change, and track which cases flip.
Run this: pick one existing agent test and run it ten times unchanged. If it doesn't pass ten out of ten, it's asserting a path, and it will either flake or get deleted.
4. Integrations break in ways the demo never sees
The demo talks to one clean API. Production talks to many. Some were built for people clicking buttons, some throttle under load, and all of them have credentials that expire.
Agents stress integrations differently from traditional code. They retry, they sometimes invent parameters, and they read error text as if it were content. A rate limit becomes a burst of retries. An expired token becomes a stream of 401 responses that the agent may "work around". Credential handling is enough of its own problem that we cover it separately in the authentication piece.
Run this: list every system your agent calls. For each, write down three things: how it fails (timeout, throttle, auth), what the agent currently does when it fails, and whether anyone has tested that path deliberately. Any integration with an untested failure path is where your first production incident is waiting.
5. Nobody owns the agent end to end
Ask who gets paged when the agent fails at 2 a.m. In many organisations the honest answer is nobody in particular. The data team built the prompts, the platform team runs the infrastructure, the business team owns the outcome, and the agent crosses all three.
An agent without an end-to-end owner fails slowly. Errors aren't triaged, the evaluation set isn't maintained, and the configuration drifts. Then a change that flips the agent from suggest to apply goes through without anyone weighing the risk. This is the same gap that makes agent governance start with an inventory: you can't govern authority you can't see or assign.
Run this: for each production agent, name one person who can switch it off, owns its evaluation set and answers for its business outcome. If the answer is a team alias or a committee, the agent has no owner.
The pre-production gate
The five patterns turn into five questions. Ask them before launch and again after any change to scope, tools or model. An agent that can't answer all five isn't ready, however good the demo was.
# | Gate question | Evidence that passes | Common failing answer |
1 | What is the end-to-end success rate of the main task, measured on real runs? | A number from sampled logs, plus the step count | "The demo works every time" |
2 | Does every claimed side effect get verified against the system of record? | A verification check per tool, and an escalation on tool failure | "The agent reports success" |
3 | Do tests assert outcomes statistically rather than paths? | Pass thresholds over repeated runs, plus an adversarial suite in CI | Single-run tests on tool order |
4 | Has every integration's failure path been tested on purpose? | Test results for timeout, throttle and expired-credential cases | "We'll handle that if it happens" |
5 | Who owns the agent end to end? | One named person with authority to stop it | A team alias |
Tradeoffs and exceptions
Narrow scope costs flexibility. A tightly scoped agent does less, and some stakeholders will see that as a step backwards from the demo. That's usually the right trade: one workflow done reliably is worth more than five done unpredictably, and scope can grow once the first workflow holds.
Verification costs latency and engineering time. Checking every side effect adds calls and code. For read-only agents that only draft text for a person to review, a lighter gate is reasonable: the human review is the verification. The full gate matters most where the agent writes to systems, moves money or contacts customers.
Statistical tests are slower and less comfortable. Repeated runs cost compute and produce a pass rate instead of a clean green tick. Teams that resist this usually end up muting flaky tests instead, which is worse.
Implementation reality
The gate only works if it's attached to something that already happens. Put it in the release checklist for agent changes. Make question 5 a required field on the agent's registration. And schedule the 50-task verification sample as a recurring job with a named owner, not a one-off audit. Expect the first run of the gate to fail on question 2 or 4; that's normal, and it's the point.
Each of these failure patterns maps to a stage that was skipped before launch. What is AI agent management? lays out those stages and the evidence each one needs before an agent earns more autonomy.
Before an agent moves from pilot to production, the Pilot-to-Production Readiness Checklist scores the five capabilities these failures exploit, plus the ownership gates and the metrics to commit to before launch.
What to do next
This week, run check 2 on your highest-volume agent: sample 50 completed tasks and verify each against the system of record. It takes an afternoon. It will tell you more about your agent's real reliability than any dashboard you have now, and it gives you a concrete number to bring to the go/no-go conversation.
Final takeaway
Agents that fail in production rarely fail mysteriously. They fail where the system lets a probabilistic step pass without checking it, and that happens in the same five places. Gate on those five, and the demo stops being the last time the agent was tested.
Sources
Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Moshkovich, D., Mulian, H., Zeltyn, S., Eder, N., Skarbovsky, I., Abitbol, R., Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems, arXiv 2503.06745, March 2025 (calculator agent, 50 examples × 5 runs: mean CV of 63% in execution-path structure, 19% in accuracy). https://arxiv.org/abs/2503.06745


