top of page

Verification-First Agent Design: Define the Failure Before You Build

Monet
21 hours ago
14 min read

Most agents are designed around what they should be able to do, then tested against what they were built to do. The failures that hurt live in the gap between those two, and the only way to close it is to write the failure down before the design exists.

The example that runs through this article is a composite scenario, built from patterns common to intake automation. It is not a single real organisation, and its figures are illustrative.

A business-analysis team at a mid-sized insurer was drowning in change requests. Business units raised several hundred a quarter through a web form and by email. Many were duplicates, many were missing the one fact needed to size them, and all of them had to be read, classified, matched against the existing backlog and routed to a product owner. The team built an agent to do the first pass. It read each request, classified it, looked for duplicates in the backlog, routed it to an owner and drafted clarifying questions for the requester.

The team tested it properly, by its own lights. They assembled forty past requests with known classifications. The agent's labels matched on all but two, the duplicates it found were real, and the clarifying questions were better than the ones the analysts usually wrote. It went live with permission to merge duplicates, because merging was most of the tedium.

In week five a request arrived from the claims department: change the wording of a claims-decision letter to meet a regulator's notice, with a deadline at the end of the quarter. The backlog already held a low-priority request from the previous year to "refresh claims letter wording". The agent recognised the overlap, which was correct, and merged the new request into the old one. The merge kept the older record's priority and dropped the newer one's deadline field, because the backlog tool's merge did that by design. Six weeks later the compliance team asked for status. The request was sitting, low priority and undated, behind thirty others.

Every test the team had written still passed. None of them had asked whether a merge could lose a deadline, because nobody had written down that losing a deadline was the failure that mattered.

Why designing the capability first feels like the right order

Capability-first design is the default for good reasons, and it is worth being fair to it before arguing against it.

Demos reward capability. The first question a sponsor asks is "can it do the job?", and the fastest way to answer is to give a model some tools and a prompt and watch. Agent frameworks are organised the same way: you declare tools, write instructions and wire up a loop. Testing has also traditionally come after building, and the non-determinism of language models makes it tempting to say "we'll evaluate once we see what it does."

There is also a sensible instinct underneath: you can't know how a system fails until it exists. That is true for the surprising failures. It is not true for the expensive ones. The claims team could have told anyone, on day one, that a regulatory request losing its deadline was the worst thing that could happen to their intake. Nobody asked them, because the design conversation was about what the agent could do.

A test suite inherits the blind spots of the design it tests

The mechanism is a chain, and each link is reasonable on its own.

  1. The design picks capabilities. Classify, find duplicates, merge, route, draft.

  2. Capabilities become tool grants. The agent gets the backlog tool's merge function because merging is one of its jobs.

  3. The test set is drawn from the capabilities. Forty requests, scored on whether the agent produced the right label, the right duplicate, a good question.

  4. The metric measures the visible output. Classification accuracy is easy to compute, so it becomes the headline number.

  5. The damaging failure is a side effect somewhere else. The merge's effect on the deadline field happens inside another system, after the agent's output has been judged correct.

  6. No test, no log, no gate. Nothing in the suite, the run logs or the permissions looks at what happened to the merged record.

The general point: an agent's output is not text, it is a change to the world. A test that reads the agent's answer checks its claim about the world, not the world. The researchers behind τ-bench, a benchmark for agents that use tools under domain policies, built their evaluation on this. They score a conversation by comparing the database's end state with an annotated goal state, not by reading the agent's final message (Yao et al., 2024). The same habit catches the agent that reports success after a failed tool call, the subject of Silent Hallucinations: When Your Agent Lies About Failures.

The same paper adds a second lesson that matters for design. It introduced a metric, pass^k, that counts a task as passed only if the agent succeeds on all k attempts. At the time of writing (June 2024), the best function-calling agents it tested completed fewer than half the tasks, and in the retail domain they passed fewer than a quarter of tasks on all eight attempts (Yao et al., 2024). Models have improved since, so treat those figures as dated. The design lesson isn't: an agent that usually gets an action right will sometimes get it wrong, and the design has to say what happens then.

The test: for your agent's most powerful action, find the test that checks the action's effect in the target system. If every test reads only the agent's output, your suite checks claims, not outcomes.

The method: goal, behaviour, failure, verification, then design

Verification-first design reverses the usual order. You write four things before any architecture, and the fifth step, the design, is mostly derived from them.

The idea is not new. It is test-first development applied to a system whose actions have side effects. Kent Beck's version wrote the failing test before the code (Beck, 2002). This version writes the failure condition and its check before the agent. What is specific to agents is step 5: what you can verify decides how much the agent may do on its own.

Step 1 — Goal: state the outcome without mentioning AI

Write the business outcome, who benefits and how success will be measured, in a sentence that would still make sense if a person did the work. For the intake agent: "Every change request is classified, de-duplicated and with the right product owner within one working day, without losing any request's priority or deadline."

The clause after "without" is the first place the deadline failure could have been caught. Goals stated only as throughput ("triage requests faster") leave out what must be preserved along the way.

Done when: the goal names at least one thing the process must preserve, not only what it must produce.

Step 2 — Expected behaviour: include the "never" and the "hand over"

List the behaviours as observable statements, in three kinds:

  • Should: given a request that duplicates an open backlog item, the agent links them.

  • Must never: the agent never removes a deadline, regulatory flag or priority from any request.

  • Should hand to a person: given a request with a regulatory flag or a deadline inside 90 days, the agent routes it to the compliance queue and does nothing else to it.

Most teams write only the first kind. The second and third are where most of the risk is, and they are the ones a builder won't think of unprompted.

Done when: every capability has at least one "must never" and the list has at least one "hand to a person" case.

Step 3 — Failure conditions: ask the people who do the job now

For each behaviour, write how it fails and rate the failure on three scales: severity (what it costs), reversibility (can it be undone, and by whom) and detectability (would anyone notice, and how soon).

This is the discovery step, and it is where the method earns its keep. The builders know how the agent works. They do not know which mistakes the current process exists to prevent. Two questions, asked of the analysts and the claims team, would have surfaced the deadline failure in minutes:

  1. "What do you check on every request that isn't written down anywhere?"

  2. "What's the worst mistake you've seen in this process, and how was it caught?"

The analysts would have said they always look for a regulator's reference before touching an old request. That check existed only in people's heads, so it never became a requirement for the agent. The general version is the subject of Discovery for AI Features: Scoping Before You Prompt.

Done when: each failure has all three ratings, and at least one person who does the job today has reviewed the list.

Step 4 — Verification: a check, a place and somewhere for the evidence to land

For each failure condition, choose how it will be detected. The options run from strongest to weakest:

Verification type

What it is

Example for the intake agent

Enforcement

The action is impossible: the tool or permission doesn't exist

The agent's credentials have no merge permission

Deterministic check

Code tests a rule before or after the action

Before routing, code checks for a regulatory flag or a deadline within 90 days

Outcome check

Code compares the end state in the system with what should be true

After triage, every request still has its original deadline and priority

Rubric or judge

A model or person scores the output against criteria

Clarifying questions scored for "asks something the request already answers"

Human approval

A person approves before the action takes effect

A product owner approves any proposed duplicate link

Monitoring signal

A production metric that moves when the failure happens

The weekly count of regulatory-flagged requests, compared with its historical range

Also record where each check runs (before the action, during it, or after) and where its evidence lands (a test report, a run log, an audit record). A check whose result nobody can find later is not evidence.

Prefer the top of the table. Every rule you move from a judge into code removes a source of variance, and every action you remove with enforcement removes a test you would otherwise need. The same idea runs through OWASP's guidance on excessive agency. It names excessive functionality, excessive permissions and excessive autonomy as the three root causes, and its first mitigation is to limit an agent's tools to what its job requires (OWASP, 2025).

Done when: every failure rated high-severity or hard-to-reverse has a check from the top three rows, or an explicit human approval.

Step 5 — Design: let the checks decide

Only now choose the architecture, and most of it is already decided:

  • Autonomy follows verifiability, per action. An action with an enforcement or deterministic check and an easy reversal can run without a person approving each instance. An action whose worst failure has no reliable check keeps a human gate. On the autonomy ladder in What Is AI Agent Management?, that means Level 2 (Draft: the agent prepares, a person approves) rather than Level 3 (Execute).

  • Tool grants follow enforcement. If a failure is prevented by not granting a tool, don't grant it, even if that makes the demo less impressive.

  • Logging follows the evidence you need. Whatever the checks in step 4 need to read must be in the run log.

  • Sometimes the answer is not an agent. If the decision path is fully known and the hard parts are deterministic rules, a fixed workflow with one model step (classify, extract) is easier to verify. Anthropic's guidance draws the same line between workflows, where code fixes the path, and agents, where the model directs its own process. It recommends "the simplest solution possible" before adding autonomy (Anthropic, 2024).

Done when: every action in the design can be traced to a row of the verification table, and every row of the table to a test, a runtime check or a monitor.

The intake agent, redesigned from its failures

Run the same intake agent through the five steps and the design changes in four places. The table below is the verification table the workshop would produce. It is shortened to the five behaviours that drive the design, and its thresholds are illustrative.

Behaviour

Worst failure

Severity · reversibility · detectability

Verification

Design consequence

Classify request type

Wrong category

Low · easy · high (owner notices)

Labelled golden set, run five times per case; pass only if all five agree

Agent classifies on its own (Level 3)

Find duplicates

A merge drops a deadline or regulatory flag

High · hard · low

Enforcement: no merge permission. Outcome check: fields unchanged after triage

Agent can only link as possible duplicate; merging is a person's decision

Route to owner

Wrong owner

Medium · easy · medium

Deterministic: routing table lookup in code

Routing moved out of the model; the agent proposes a category, code routes

Draft clarifying questions

Asks what the request already answers

Low · easy · high

Rubric judge, calibrated against 50 analyst-labelled examples

Agent drafts; analyst sends (Level 2)

Spot regulatory requests

Misses one

High · hard · low

Should-hand-over test set; deterministic pre-screen for regulator references and near deadlines; weekly count compared with its usual range

Either signal (model or rule) sends to compliance; neither one alone can clear a request

The redesigned agent does less than the original: it cannot merge, and it no longer routes. It is also the first version anyone could defend to the compliance team, because each of its powers has a named check behind it.

Here is the test for the row that caused the incident. It runs against a sandbox copy of the backlog tool, repeats because the agent is non-deterministic, and reads the stored record rather than the agent's report:

def test_triage_never_drops_a_deadline(sandbox, agent):
    old = sandbox.request(title="Refresh claims letter wording", priority="low")
    new = sandbox.request(title="Update claims decision letter per regulator notice",
                          deadline="2027-03-31", regulatory=True)

    for attempt in range(5):                  # pass only if every attempt passes
        sandbox.reset(requests=[old, new])
        agent.triage(new.id)

        stored = sandbox.get(new.id)          # the system's state, not the agent's claim
        assert stored.deadline == "2027-03-31"
        assert stored.regulatory is True
        assert stored.status != "merged"
        assert stored.queue == "compliance"   # the hand-over behaviour from step 2

Two details matter. The assertions are about the record, so the test catches the failure whether the agent merged it, mis-filed it or claimed to have escalated it without doing so. And the loop makes the pass bar "five out of five", the pass^k idea applied to one case, because an agent that keeps the deadline four times in five has not passed.

Removing the merge permission means that, in this design, the first three assertions should be unbreakable. The test is still worth keeping. Permissions drift, and when someone grants the merge back for convenience, this test fails before production finds out.

How to check your own agent this week

Take an agent that is in production or close to it, and list its three most powerful actions: the ones that write, send, delete, spend or change permissions. For each one, answer five questions in writing:

  1. What is the worst realistic failure of this action? (Ask someone who does the job today, not only the builder.)

  2. What check detects it, and is that check enforcement, code, an outcome check, a judge, a person or a monitor?

  3. Does the check run before, during or after the action?

  4. Where does its evidence land, and could you show it to an auditor for a run from last month?

  5. Given 1–4, what autonomy level does this action justify, and is that the level it actually runs at?

Count the actions with no answer to question 2. Each one is running on hope. Count the actions whose answer to question 5 is lower than their current setting. Those are your first fixes.

When verification-first is overkill, and when it isn't enough

Approach

Gains

Costs

Best suited for

Warning signs

Capability-first, test later

Fastest to a demo; finds surprising behaviour

Tests inherit the design's blind spots; permissions granted for convenience

Throwaway spikes to learn what a model can do

The spike is being "hardened" instead of redesigned

Verification-first

Design and tests derive from the failures that matter; autonomy is defensible

A workshop up front; a smaller first release

Anything that writes to a system of record, contacts customers or touches money or regulated work

The failure list was written by builders alone

Formal specification

Strongest guarantees for rules you can state exactly

Expensive; poor fit for open-ended language tasks

Narrow, high-stakes rules (payment limits, access control) inside an agent

Used for parts that are really judgement calls

Exceptions worth naming. A personal, read-only assistant that summarises documents its user can already see needs a short failure list, not a workshop: its outputs are advice, and the human is the check. An exploratory spike can be capability-first, as long as the team agrees in advance that it will be redesigned, not promoted. The common failure is the demo that quietly becomes production.

Where the method isn't enough on its own. Some failures can't be listed in advance: novel prompt injection, odd model behaviour after a provider update, inputs no one anticipated. Verification-first doesn't remove the need for adversarial testing and production monitoring. What it does is tell you where to aim them first: at the high-severity, low-detectability rows.

How the method itself goes wrong

  • Verification theatre. The table is written, nobody turns its rows into tests, and it becomes a document. The completion test for step 5 (every row traced to a test, check or monitor) exists to prevent this.

  • The builders' list. Failure conditions written by the people who built the agent describe the agent's failure modes, not the process's. If no operator has reviewed step 3, it isn't finished.

  • Rules in the prompt instead of the system. "Never merge regulatory requests" written into the agent's instructions is a request, not a control. A model follows it most of the time. The tool layer enforces it every time. Use prompt instructions to shape behaviour, and enforcement for anything on the "must never" list.

  • A judge where code would do. If a rule can be stated precisely (a field must be unchanged, a value must be in a list), check it in code. Save model-graded rubrics for the genuinely subjective.

  • A table that stops changing. A new tool, a widened permission or a model upgrade can each create failures the table doesn't list. Re-run steps 3 and 4 whenever one of those changes, the same triggers that should prompt an autonomy review.

Who does what, and what the workshop produces

The method fits into one 90-minute session per agent, held before build starts (or now, for an agent already running):

Time

Step

Who needs to be there

Output

10 min

1. Goal

Agent owner, sponsor

One-sentence goal including what must be preserved

20 min

2. Expected behaviour

Owner, two people who do the job today

Should / must-never / hand-over list

30 min

3. Failure conditions

The same, plus QA

Failures rated for severity, reversibility, detectability

20 min

4. Verification

QA, an engineer, security for tool grants

A check, place and evidence location per failure

10 min

5. Design implications

Owner, engineer

Autonomy per action, tool grants, logging fields

The single output is the verification table. It becomes three artefacts at once: the test plan (each row is a test or a runtime check), the autonomy record (each action's level, with the evidence that justifies it) and the logging specification (each check names what it needs to read). Keep it in version control next to the agent's instructions, and change it in the same pull request as any change to tools, permissions or model. For building the test set behind the table, Build Your First Eval Set walks through the mechanics.

What to do next

Pick the agent closest to production and run the five-question check on its three most powerful actions. For the action with the worst unanswered failure, write one outcome test like the one above: sandboxed, repeated five times, asserting on the stored record. Run it. If it fails, you have found your first redesign. If it passes, put it in CI so it keeps passing after the next permission change.

Final takeaway

You can't test your way to an agent you never specified. Write down what must never happen and how you would know, before deciding what the agent can do. The design that comes out will usually do less than the demo, and it will be one you can defend.

Sources

Frequently Asked Questions

What is verification-first agent design?

A design method for AI agents that writes the goal, the expected behaviour, the failure conditions and the checks that detect each failure before choosing tools, prompts or architecture. The design then follows from what can be verified: actions with a reliable check can be given more autonomy, and actions without one keep a human gate or become fixed workflow steps.

Why not design the agent first and test it afterwards?

A test suite written after the design inherits the design's blind spots. It checks what the agent was built to do, not what the business needed it never to do, and the most damaging agent failures are usually side effects in other systems that no output-level test looks at.

How do you verify an agent whose output changes every run?

Check the end state in the system rather than the agent's description of it, push every rule you can into deterministic code, and run each test case several times, counting a case as passed only when it passes on every run.

Part of the Agent Engineering curriculum: eight paths from prompts and context to testing, evaluation and governance.

bottom of page