Agent, Workflow, Automation or RPA? Choose the Least Autonomous Thing That Works
Autonomy is not a feature you add for free. Every decision you hand to a model is a decision your tests now have to cover, so the right design is the one with the least autonomy that still does the job.
The example in this article is a composite scenario, built from patterns common to finance operations. It is not a single real organisation, and its figures are illustrative.
The accounts-payable team at a mid-sized distributor had a sponsor's brief: "build an AI agent to resolve invoice exceptions". The team handled about 1,200 exceptions a month, the backlog was growing, and the vendor demo had shown an agent reading an invoice, opening the ERP, checking the purchase order, emailing the supplier and closing the case. Everyone liked the demo. The quality lead was asked to start writing tests.
Before writing any, she pulled a month of exceptions from the ERP and grouped them by reason code. The picture did not look like the demo:
Exception type | Share of volume | What the team actually does |
Price or quantity variance inside an agreed tolerance | 55% | Approves it. Every time. |
Goods receipt missing | 20% | Looks up the delivery in the warehouse system, then matches or chases |
Supplier credit notes and remittances arriving as PDF attachments | 15% | Reads the PDF, keys the fields, matches them to the invoice |
Supplier disputes | 8% | Reads the history across email and the ERP, decides, replies |
Everything else (duplicates across legal entities, odd one-offs) | 2% | Investigates across several systems |
More than half the volume was a rule the ERP could already apply, if anyone had configured the tolerance. Another fifth needed one lookup in a warehouse system that had no API. The PDFs needed reading, but what to do with the fields afterwards was fixed. The disputes needed a person's judgement. Only the last 2%, open-ended investigations whose steps couldn't be listed in advance, looked like work for an agent.
The project shipped as a tolerance setting, a screen bot, an extraction step inside a fixed workflow, and an assistant that drafts dispute summaries. The "agent" became a read-only investigator for the remaining 2%. It was cheaper to build, and it was far cheaper to test.
Why "build an agent" is the default ask
It is easy to see how the brief got written. Agents are what the market is selling, and a demo of one doing a whole job end to end is far more persuasive than a demo of a tolerance setting. "Agent" has also become a loose label: vendors apply it to chatbots, to workflows with a model step, and to scripts with a language-model call inside. A sponsor who asks for an agent usually means "make this work go away", not "give a model the authority to choose its own actions".
There is a real reason underneath, too. Exceptions are messy, and messy work looks like it needs judgement. Sometimes it does. The mistake is to treat the whole category as messy when the mess is concentrated in a small part of it.
Autonomy is paid for in test surface
The cost of autonomy is easiest to see from the testing side, which is where the quality lead started.
A workflow with a model step has a bounded output. If the model classifies an exception into one of five reasons, the test question is "is the label right?", and a labelled sample answers it. The code around the model is deterministic and is tested the ordinary way.
An agent's output is a sequence of actions, and the space of sequences grows fast. As an illustrative calculation: an agent with five tools that may take up to six steps has 5⁶ = 15,625 possible ordered sequences of exactly six calls, before counting the arguments to each call. You can't test them all, so you test outcomes and boundaries instead (the subject of How to Test an AI Agent). Every degree of freedom you hand to the model still widens what can go wrong between the input and the outcome.
OWASP names this directly. Excessive autonomy is one of the three root causes of its "excessive agency" risk for LLM applications, alongside excessive functionality and excessive permissions (OWASP, 2025). And the most widely cited practitioner guidance on agents makes the same recommendation from the build side: find "the simplest solution possible", and only increase complexity when needed (Anthropic, 2024).
The test: for the work you've been asked to automate, write down how many distinct actions the proposed agent could take and how many steps it may chain. If you can't say, nobody can test it yet.
Five designs, from least to most autonomy
The vocabulary is muddled, so here are working definitions. They describe who decides what happens next.
Design | Who decides the next step | Where a model fits | Example from the case |
Deterministic automation | Code, from rules written in advance | Nowhere | Auto-approve variances inside tolerance |
RPA | Code; a bot operates a user interface the way a person would | Nowhere (the bot follows a script) | Read delivery status from the warehouse system's screens |
Workflow with a model step | Code fixes the path; the model fills one bounded step | Classify, extract, summarise | Extract credit-note fields from a PDF, then validate them in code |
Assistant | A person, on every case | Suggests, drafts, summarises | Draft a dispute summary from the email and ERP history |
Agent | The model, within limits you set | Chooses tools and steps until the goal is met | Investigate a suspected cross-entity duplicate |
Anthropic's distinction is the useful line through the middle of this table. Workflows are systems where models and tools are orchestrated through predefined code paths. Agents are systems where the model dynamically directs its own process and tool use (Anthropic, 2024). Everything above the last row keeps the path in code. That is what makes those designs testable with ordinary tools.
The decision tree
Ask the questions in order, for each type of work, not for the whole process. Stop at the first yes.
Is every step known in advance, with structured, stable inputs? Use deterministic automation. Most "AI" projects contain a large slice of this, often a rule that already exists in someone's head or in an unused configuration setting.
Is the only obstacle a system you can reach only through its screens? Use RPA for that step, and keep the logic in code. RPA breaks when screens change, so prefer an API wherever one exists.
Is the uncertainty only in reading messy input? Classifying, extracting and summarising are bounded jobs. Use a workflow with one model step, and validate the model's output in code before anything acts on it.
Does a person decide and act on every case? Then you're building an assistant. Its quality question is whether people trust the drafts appropriately, so watch for edits dropping to zero, which can mean rubber-stamping rather than accuracy.
Can each action be constrained, checked and reversed? Then an agent is a reasonable design, at the lowest autonomy level the evidence supports. If not, narrow the scope until each action can be checked, or keep the human.
Question 5 is the gate that matters most, and it is not about the model's ability. A highly capable model taking an unverifiable, irreversible action is a worse design than a weaker model taking checkable ones. How to answer it with evidence is set out in Verification-First Agent Design.
How the invoice project changed shape
Running each exception type through the tree produced a design with five parts, not one:
Exception type | Answer | Design | How it's verified |
Variance inside tolerance (55%) | Q1: yes | ERP tolerance rule, configured and documented | Unit test on the rule; monthly sample of auto-approvals |
Goods receipt missing (20%) | Q2: yes, for the lookup | RPA reads delivery status; code matches or sends a templated chase | Bot run log; matched-vs-chased counts reconciled daily |
Credit notes in PDFs (15%) | Q3: yes | Model extracts fields; code validates totals and PO numbers before posting | Labelled extraction set; any validation failure goes to a person |
Disputes (8%) | Q4: yes | Assistant drafts the history summary and a reply; the clerk decides | Clerks' edit rate on drafts, tracked weekly |
Cross-entity investigations (2%) | Q5: yes, read-only | Agent with read access to three systems; produces a findings note | Findings checked against the clerk's conclusion on every case |
The agent that survived is deliberately weak. It can read, search and write a note. It cannot post, pay or email. On the autonomy ladder in What Is AI Agent Management?, that is Level 1, Assist. It earns more only if its findings keep matching the clerks' conclusions.
The decisive move was not technical. It was the quality lead pulling a month of reason codes before anyone wrote a prompt. A data question ("what is this work, by type and volume?") answered a design question that a demo could not. The same move is the first stage of Discovery for AI Features: Scoping Before You Prompt.
The agent-or-not worksheet
Use this for each type of work in the process you've been asked to automate. Each question asks for evidence, not opinion.
Question | Evidence that answers it | A "yes" looks like |
1. Known steps, structured inputs? | A written procedure; the fields the work uses; how often inputs change format | The procedure fits on a page and nobody deviates from it |
2. Screen-only system? | Whether the target system has an API or export; how often its screens change | No API, and screens stable for a year |
3. Uncertainty only in reading input? | Examples of the input; what happens after it is read | After reading, the next steps are always the same |
4. Person decides every case? | Who signs off today, and whether that can change | Policy or risk requires a named person to decide |
5. Actions constrained, checked, reversible? | For each action: the check that detects its worst failure; how to undo it | Every action has a check that reads the system, and a tested undo |
Volume and value | A month of real cases, grouped by type | You know which type is most of the work |
If you can't fill in the evidence column for a type of work, that's a discovery task, not a build task.
When an agent really is the right answer
The tree shouldn't be read as a bias against agents. Some work genuinely needs one, and the signs are consistent:
The path depends on what you find. Investigations, research across many sources and diagnosis all have next steps that depend on the last result.
The check is cheap even when the path isn't. Coding agents are the clearest case. The steps to fix a bug are unpredictable, but a test suite in a sandbox checks the result every time. Anthropic's guidance recommends exactly that combination: extensive testing in sandboxed environments, with guardrails (Anthropic, 2024).
The actions are reversible or read-only. An agent that only reads and reports can run at low risk while it earns evidence.
The design options aren't exclusive. A workflow can call an agent for one bounded step, as the invoice process did for the 2% of cases, and an agent can use RPA or a workflow as one of its tools.
Design | Gains | Costs | Best suited for | Warning signs |
Deterministic automation | Cheap, fast, fully testable | Breaks on anything unanticipated | High-volume, rule-shaped work | Exceptions piling up in a manual queue |
RPA | Works with systems that have no API | Brittle when screens change; hard to version | A single legacy step inside a larger flow | Bots maintained more than they run |
Workflow + model step | Model handles messy input; path stays testable | Needs validation code around the model | Classification, extraction, summarisation | Model output used without a check |
Assistant | Human judgement stays in the loop | Throughput limited by people | Decisions with policy, money or relationships at stake | Edit rate near zero (rubber-stamping) |
Agent | Handles work whose path can't be scripted | Large test surface; cost and latency per task; needs ongoing evidence | Investigation, research, sandboxed coding | Granted write access before read-only results were checked |
How this decision goes wrong
Deciding for the whole process. One label for the whole process hides the mix. Classify by type of work, with volumes.
Letting the demo set the scope. A demo shows the case it was built for. Ask what share of real volume that case represents.
Agent-washing. A vendor's "agent" may be a workflow, which is good news for testing. Ask who decides the next step: code or the model.
RPA as a permanent answer. Screen bots are a bridge to an API, not a destination. Put a review date on every bot.
Promoting an assistant without evidence. Moving from "the clerk sends it" to "the agent sends it" is a change of design, not a setting. It needs the evidence question 5 asks for.
What to do next
Take the process you've been asked to "put an agent on". Pull a month of real cases, group them by type and run each type through the five questions with the worksheet. Bring the resulting table, not a prompt, to the next design meeting. If one type of work genuinely reaches question 5, run Verification-First Agent Design on that type alone.
Final takeaway
The question is not whether a model could do the work. It's which design does it with the least autonomy you then have to verify. In most processes, that answer is mostly rules, a little model and a small, well-checked agent, if any.
Sources
Anthropic (Schluntz, E. and Zhang, B.), Building effective agents, 19 December 2024. Definitions of workflows and agents; "the simplest solution possible"; testing in sandboxed environments.
OWASP GenAI Security Project, LLM06:2025 Excessive Agency, 2025. Excessive functionality, permissions and autonomy.
Frequently Asked Questions
What is the difference between an AI agent and a workflow?
In a workflow, code fixes the sequence of steps and a model, if there is one, fills in a bounded part of it, such as classifying a document or extracting fields. In an agent, the model decides which steps to take and which tools to call, and keeps going until it judges the goal is met.
What is the difference between an AI agent and RPA?
Robotic process automation drives a system's user interface with fixed rules, usually because the system has no API. It doesn't interpret or decide. An AI agent chooses its own actions. RPA can be one of the tools a workflow or an agent uses.
When should you use an AI agent instead of automation?
When the path to the outcome can't be known in advance and each action the agent could take can be constrained, checked and reversed. If the path is known, a workflow is cheaper to run and far easier to test. If the actions can't be checked, keep a person in the decision.
Part of the Agent Engineering curriculum: eight paths from prompts and context to testing, evaluation and governance.


