Where Human Approval Belongs: the Autonomy Ladder Per Action
An agent doesn't have one autonomy level. Each action it can take has its own, and a person's approval belongs only on the actions where no check can do the job.
The example in this article is a composite scenario built from patterns common to finance operations. It is not a single real organisation, and its figures are illustrative. The approval logs are simulated from a stated model of reviewer behaviour, not observed. The code is real: every output block below was produced by running it.
Six weeks after the duplicate-invoice investigator at a mid-sized distributor was allowed to draft reversing journal entries, the accounts-payable manager sent the finance director a one-page proposal. The clerks had approved 99.0% of the agent's drafts unchanged, in a median of 10 seconds each. The approval step was costing three clerks an hour a day and changing nothing. Promote the agent to Level 3 and let it act.
The quality lead read the same two numbers and reached the opposite conclusion. A 10-second median meant nobody was reading the drafts, so a 99% approval rate described the clerks, not the agent. She had a third number. After the double reversal told in Designing Tool Contracts and Least Privilege for AI Agents, which two clerks had approved, she had started planting known-bad drafts in the queue: a reversal for the wrong entity, an amount that included tax the original didn't, a duplicate already reversed the week before. Over the six weeks she planted twelve. The clerks caught three.
So the manager was right that the approval step wasn't working, and the quality lead was right that the agent hadn't earned Level 3. Both of them were treating the agent as a single thing with a single level, and that was the real mistake.
Why promoting the whole agent feels right
The autonomy ladder in What Is AI Agent Management? runs from Level 0, Manual, to Level 5, Optimize. At Level 2, Draft, the agent prepares the action and humans approve each one before it executes. At Level 3, Execute, it performs approved low-risk actions after deterministic checks while humans handle exceptions and sample completed work. The evidence needed to move to Level 3 is "a supervised period at Level 2 with a low, stable override rate; reversal tested; limits in place".
Read quickly, that sounds like a promotion for a person: one role, one level. The approval rate is the obvious evidence because the queue produces it for free. The anchor post names the trap: the reviewer change rate tells you whether the agent is ready, and approval time tells you whether the review is real. When approvals take seconds, the first number can't be interpreted.
The proposal also bundled actions with nothing in common: finding invoices, writing a findings note, putting an invoice on hold, drafting a reversal and, through the posting service, posting one. Putting an invoice on hold is undone with one click. A posted entry in a closed period is corrected only with another entry and an explanation to the auditors.
The test: before an approval rate is used as evidence for promotion, put the median approval time next to the time a genuine review of one item takes. Time ten real reviews to get that figure. If the median is under a quarter of it, the approval rate is evidence about the reviewers.
Approval everywhere stops being a check
The clerks weren't lazy. The design asked them to do something people reliably do badly. Parasuraman and Manzey's review of the research on complacency and automation bias found that automation bias produces both omission and commission errors when decision aids are imperfect, occurs in naive and expert participants alike, "cannot be prevented by training or instructions", and that complacency appears under multiple-task load, when manual work competes for attention with monitoring the automation (Parasuraman & Manzey, 2010). The EU AI Act writes the same risk into law for high-risk systems: Article 14(4)(b) requires that the people assigned to oversight can "remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)" (Regulation (EU) 2024/1689).
Two features of the design made it worse. The first was arithmetic. With every action that changed anything gated per item, the queue needed about 280 reviewer-minutes a day at the pace a real check takes. Three clerks gave it 180. A queue that can't be cleared at full attention gets cleared at partial attention, and nobody records that decision.
The second was the screen. The approval view showed the drafted entry: account, amount, entity, a one-line rationale. The original invoice and the suspected duplicate were a click away. The detector below reads the approval log and checks five signals against thresholds set in advance. On the six-week log it shows the clerks opened the evidence on 15% of items:
before.csv: 720 decisions
WARN median seconds 10s vs floor 60s (25% of a genuine review)
WARN evidence viewed 15% vs 80%
WARN approved unchanged 99.0% vs 97% (high is fine only if the next two pass)
FAIL seeded caught 3/12 vs 80%
ok re-review disagree 0/60 vs 2%
clerk_a: 226 items, median 10s, evidence viewed 13%
clerk_b: 235 items, median 10s, evidence viewed 17%
clerk_c: 247 items, median 10s, evidence viewed 16%
verdict: FAIL
The last signal deserves a second look. A senior clerk independently re-reviewed 60 approved drafts and disagreed with none. With a defect rate of a few per cent, a 60-item sample will often contain zero misses, so a clean re-review proves little. Planted items give you a known denominator. You know exactly how many errors were in the queue, so the catch rate is a measurement rather than a guess. That's why the seeded check is the one that fails the gate.
The discovery move that would have prevented all this is one intake question nobody asked when the gate was designed: what does a reviewer have to look at to catch the error this gate exists for, and how long does that take? The answer (both invoices, the purchase order and any earlier reversal, about four minutes) would have shown that the screen hid the evidence and that the queue couldn't be cleared at that pace.
The test: export one week of your approval log with timestamps and decisions. Compute daily volume times genuine review minutes and compare it with the reviewer time actually available. If the load exceeds capacity, you already have a rubber stamp; the log will show it.
Score each action, not the agent
The way out is to give each action its own ceiling on the ladder, decided by four attributes of the action itself:
Consequence: the worst realistic result of one bad instance, in money, customers or obligations.
Reversibility: whether it can be undone cheaply inside the company, and whether that undo has actually been run.
Detectability: whether a check you can actually run exists, and whether it runs before the action or reliably after it.
Volume: how often it happens, which decides whether per-item review can be done properly at all.
The rule that falls out is narrow. A person approves an action when it can't be undone cheaply, when no runnable check exists so a person's judgement is the check, or when one bad instance is high-consequence and the evidence for more autonomy isn't there yet. Everywhere else, a check before the action and a sample of completed work do the job better, because a check doesn't get tired at quarter-end. OWASP frames the failure from the other side: excessive autonomy is when an application "fails to independently verify and approve high-impact actions", and its remedy is to "require a human to approve high-impact actions before they are taken" (OWASP, 2025). High-impact actions, not all actions. NIST's AI RMF makes the same distinction for whole systems: "Some AI systems may not require human oversight… Other systems may specifically require human oversight" (NIST, 2023, Appendix C). Applied per action, the line runs through a single agent.
A small scorer turns those four attributes into a ceiling and a gate. The rules are applied in order and the first match decides. This is the investigator, scored action by action:
PER-ACTION CEILINGS
action L gate review min/day reason
find_and_match_invoices 1 NONE (log) 0 no side effect; reading and recommending is Assist
write_findings_note 1 NONE (log) 0 no side effect; reading and recommending is Assist
place_invoice_on_hold 3 AUTO + sample 10% 5 pre-check runs before effect; undo tested
draft_reversal_upto_5000 3 AUTO + sample 10% 8 pre-check runs before effect; undo tested
draft_reversal_over_5000 2 APPROVE 24 high consequence; stays gated until evidence says otherwise
post_journal_entry 2 APPROVE 96 high consequence; stays gated until evidence says otherwise
email_supplier 2 APPROVE 30 cannot be undone once it leaves
release_payment 0 NOT GRANTED 0 not on the role card; no tool, no credential scope
per-agent Level 2 (approve every effect): 280 of 180 reviewer-minutes/day OVER CAPACITY
per-agent Level 3 exceeds the ceiling of: draft_reversal_over_5000, post_journal_entry, email_supplier
Walk the two rows that surprised the AP manager. Putting an invoice on hold reached Level 3 because a deterministic check exists (the ledger match on supplier, amount and purchase order) and the undo has been rehearsed: release the hold, and the payment run picks the invoice up again. Drafting a reversal under 5,000 reached Level 3 for a less obvious reason. A draft posts nothing, and posting keeps its own gate. Approving the draft and the posting had put two people in front of the same decision, and each assumed the other was the real check. Removing the first approval concentrates attention on the one that matters.
The supplier email stays at Level 2 for a different reason: what the agent tells a supplier about a suspected duplicate can't be checked by code and can't be recalled once sent. Releasing payment isn't on the role card at all, so it sits at Level 0 with no tool and no scope. The scorer never assigns Level 5. Optimize means adjusting its own approach within fixed limits, which is a decision about the whole agent's objectives, not about whether one action needs a person.
The per-action design needs 163 reviewer-minutes a day against 180 available. That margin is what lets the remaining reviews be real.
The test: list every action your agent can take, including the ones it performs through another service. Score each on the four attributes, then compare each one's gate with what happens today. Every approval that none of the three reasons put there can be replaced by a check. Every action without a check that runs unapproved needs one.
Design the gate so the reviewer has something to check
Moving approval to the right actions is half the work. The other half is making each remaining approval a check rather than a ritual. The distributor rebuilt the posting gate around four changes:
Evidence on the screen, not behind a click. Both invoices side by side, the purchase order and delivery notes, the ledger match result, and every earlier reversal for that supplier in all three entities.
One specific question instead of "Approve?". The double reversal showed which error mattered most, so the gate asks it: is there already a reversal for this invoice in any entity? A reviewer can answer that from the screen in a minute or two, and can't honestly answer it without looking.
Rules about who and when. The approver is never the agent and never whoever drafted the entry. A timeout escalates to the AP lead after four working hours; it never auto-approves.
Planted items, permanently. Two known-bad drafts a week, tagged in a table reviewers can't see and pulled before anything posts, so the catch rate is always current.
Anthropic's guidance on agents describes pausing "for human feedback at checkpoints or when encountering blockers" (Anthropic, 2024). The redesign decides what that checkpoint shows and asks. Two weeks of the new gate (simulated from the same behaviour model, with evidence now on screen):
after.csv: 200 decisions
ok median seconds 96s vs floor 60s (25% of a genuine review)
ok evidence viewed 95% vs 80%
ok approved unchanged 95.8% vs 97% (high is fine only if the next two pass)
ok seeded caught 8/10 vs 80%
ok re-review disagree 0/40 vs 2%
clerk_a: 65 items, median 90s, evidence viewed 95%
clerk_b: 63 items, median 110s, evidence viewed 98%
clerk_c: 62 items, median 91s, evidence viewed 92%
verdict: PASS
Read 8 of 10 with care: it sits exactly on the bar, and ten planted items is a small sample, which is why the planting continues every week. Note also what the detector does not fail on. A high approval rate is only a warning, because a genuinely good agent should be approved most of the time. It becomes a finding when the planted items or the re-review say the approvals aren't discriminating.
The test: write down, for each approval gate you run, the one question the reviewer must answer and the evidence they need to answer it. If you can't name the question, the gate is a signature, not a check.
Escalation is the agent's half of the gate
A gate covers what the agent proposes. Escalation covers what it should never have proposed: the cases where it should stop and hand over instead of drafting with confidence. For the investigator, the hand-over conditions are concrete enough to test:
a pre-check fails or can't run (the ledger is unavailable, a match field is empty);
the input is outside what it was tested on: a new entity, a currency it hasn't seen, a credit note instead of an invoice;
the evidence conflicts: the purchase orders match, the delivery notes don't;
the amount is over its limit, or an earlier action already exists on the same case.
An escalation should carry the same packet a reviewer gets: the evidence, the specific doubt, and what the agent would have done. Then it can be measured like any other behaviour: write cases where escalating is correct and cases where it isn't, and track how many should-escalate cases it catches and how many of its escalations were needed, as How to Test an AI Agent sets out for trajectory tests. Both rates matter, because an agent that over-escalates teaches the clerks to skim escalations too.
The test: write ten cases from last quarter where a clerk would have stopped and asked someone. Run the agent on them. Any case it drafted instead of escalating is a missing hand-over rule.
The approval gate worksheet
This is the T12 worksheet, one block per action, filled in for the posting action. The thresholds were set before anyone looked at the data, which is the discipline the anchor post asks for: decide what "low and stable" means in advance, so the numbers can't be argued into fitting.
# T12 Approval Gate Design worksheet. One block per action the agent can take.
action: post_journal_entry
side_effect_class: irreversible # read | draft | write | irreversible (from the tool contract)
consequence: high # worst realistic result of ONE bad action, in money or customers
reversal: costly # tested | costly | impossible. When was the undo last run?
check_you_can_run: pre # pre | post | none. Name it, or write "none"
check_named: ledger match on supplier, amount and PO; no prior reversal in any entity
volume_per_day: 24
ceiling: 2 # from the five questions, in order
gate:
approver: AP clerk; never the agent, never the person who drafted it
reviewer_sees: [original invoice, suspected duplicate, PO and delivery notes, ledger match, prior reversals]
the_one_question: Is there already a reversal for this invoice in any entity?
genuine_review_minutes: 4 # timed on ten real items, not guessed
on_timeout: escalate to the AP lead after 4 working hours; never auto-approve
measure: # thresholds set before looking at the data
seeded_items: 2 per week, known-bad, pulled before anything posts
seeded_caught: ">= 80%"
median_seconds: ">= 60 (25% of a genuine review)"
evidence_viewed: ">= 80%"
rereview_disagree: "<= 2% on 40 approved items a month"
escalate_when: # the agent's half: it stops and hands over
- a pre-check fails or cannot run
- entity, currency or document type it has not been tested on
- evidence conflicts (PO matches, delivery notes differ)
- amount over its limit, or a prior action already exists on the case
promote_when: 8 weeks at ceiling with every measure passing; reversal rehearsed
demote_when: seeded_caught below 80% two weeks running; any duplicate posting; model or prompt change
The side-effect class comes straight from the tool contract, and the gate is enforced at the tool, where the contract post put it. The worksheet adds what the contract can't: who looks, at what, for how long, and how you'll know they still are.
How it ended. The director approved per-action ceilings rather than a promotion: holds and drafts up to 5,000 moved to Level 3 for one legal entity first, with 10% sampled review. Larger drafts, posting and supplier emails stayed at Level 2 behind the redesigned gate. The demotion triggers in the worksheet were signed at the same meeting.
Tradeoffs and exceptions
Option | Gains | Costs | Use when |
Per-item approval | A person sees every instance | Reviewer load; automation bias at volume | Irreversible, uncheckable or high-consequence actions at volumes people can review properly |
Check plus sampled review | Consistent; scales; doesn't tire | Building and maintaining the check; misses what the check doesn't test | Reversible actions with a deterministic check |
Approval only at the irreversible step | One gate gets full attention | Upstream errors travel further before they're seen | Draft-then-post chains |
Planted items | A known denominator for catch rate | Must be built and pulled before effect; reviewers may learn to spot them | Every gate that runs more than a few items a week |
Exceptions. Where a law or regulator requires a person to review each decision, the per-item gate stays whatever the scorer says. Whether a system falls under Article 14 of the AI Act is a legal classification, made per system. A brand-new agent with no evidence can reasonably start with every action at Level 2. That's a starting state with exit criteria, not a design. And a "check" that is itself a model, such as an LLM judge, doesn't count as a check you can run until it has been calibrated against people on your task, as Evaluating AI Agents shows. An uncalibrated judge in that seat is another rubber stamp, just a faster one.
How approval gates go wrong
Approval as liability transfer. A gate that exists so someone can be blamed is never designed so they can succeed. Ask what the reviewer can see before asking who signs.
Auto-approve on timeout. A gate that approves when nobody answers has quietly become Level 3, set by whoever configured the timeout.
Measuring approval rate alone. It rises when the agent improves and when reviewers stop looking, and the two are indistinguishable without planted items.
Planted items that leak. A seeded draft that posts is a real error you created. Pull them in code before the effect, and test that you do.
Ceilings that only go up. The worksheet's demotion triggers are the half that gets skipped. A model change or a falling catch rate should move an action down a level automatically, pending review.
This is the operating half of Verification-First Agent Design: where a check you can run exists, design follows from it; where none exists, the action keeps a human gate.
What to do next
Pick the agent whose approval queue is busiest. Pull one week of its approval log and compute three numbers: median time-to-approve against the timed length of a genuine review, the share of items where the reviewer opened the evidence, and daily review load against reviewer time available. Plant five known-bad items over the next week and count the catches. Then score each of the agent's actions on the four attributes and fill in the worksheet for the highest-volume action that stays at Level 2.
Final takeaway
Human approval is a control with a capacity, and it fails quietly when it is spread over every action an agent takes. Give each action its own ceiling on the ladder, keep people on the actions that can't be undone or checked, and give them the evidence and one question to answer. Then keep planting errors, because a gate that has stopped catching them looks the same as one that still does until you count.
Sources
R. Parasuraman and D. H. Manzey, "Complacency and Bias in Human Use of Automation: An Attentional Integration", Human Factors 52(3), 381–410, 2010. doi:10.1177/0018720810376055. Abstract verified via PubMed (PMID 21077562): automation bias produces omission and commission errors with imperfect aids; occurs in naive and expert participants; cannot be prevented by training or instructions; complacency under multiple-task load.
Regulation (EU) 2024/1689 (AI Act), Article 14, human oversight, paragraph 4(b) on automation bias. Applies to high-risk AI systems.
OWASP GenAI Security Project, LLM06:2025 Excessive Agency, 2025. Excessive autonomy; human approval for high-impact actions.
NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023. Appendix C on human-AI configurations and oversight; MAP 3.5 on defining human oversight processes.
Anthropic, Building effective agents, 19 December 2024. Pausing for human feedback at checkpoints or blockers; stopping conditions.
Frequently Asked Questions
Where should human approval go in an AI agent workflow?
On the actions that cannot be undone cheaply, the actions where no check you can actually run exists, and high-consequence actions until evidence supports more autonomy. Everywhere else, a deterministic check before the action plus sampled review of completed work does the job better, because a check does not get tired or hurried.
What does it mean to apply the autonomy ladder per action?
Instead of giving an agent one level, give each action it can take its own ceiling, based on consequence, reversibility, whether a check exists and volume. One agent can put an invoice on hold at Level 3 while posting a journal entry stays at Level 2 with a person approving each one.
How do you tell if human approvals have become rubber-stamping?
Plant known-bad items in the queue and count how many reviewers catch them. Add median time-to-approve against how long a genuine review takes, the share of items where the evidence was actually viewed, and a sampled independent re-review. A near-100% approval rate on its own tells you nothing.
Part of the Agent Engineering curriculum: eight paths from prompts and context to testing, evaluation and governance.


