top of page

AI Agent Engineering: Build, Test and Govern Agents That Hold Up in Production

Monet
21 hours ago
7 min read

An agent is not a prompt with ambitions. It is a system that takes actions in other systems, and engineering one means being able to say, for every action it takes, how you would know it was wrong.

Consider a situation many quality leads are now in. An agent has been built to handle part of an operational process: reading requests, updating records, drafting replies. The demo went well and the sponsor wants it live next month. The quality lead is asked to "test it", and finds there is no specification to test against. There is a prompt, a list of tools and a handful of examples that worked. What the agent must never do, which of its actions can be undone, and what evidence would show it is behaving: none of that was written down, because the design started from what the agent could do.

That is the gap this curriculum is built around. There is no shortage of material on how to wire a model to some tools. There is far less on how to decide what an agent should be allowed to do, how to prove it is doing it, and how to keep proving it after the model, the prompt or the permissions change. ShiftQuality's agent engineering paths cover both halves, and they put the second half first.

What agent engineering covers

Anthropic's widely cited guide draws a distinction worth adopting. Workflows are systems where language models and tools are orchestrated through predefined code paths. Agents are systems where the model dynamically directs its own process and tool use (Anthropic, 2024). Agent engineering is the work of building the second kind safely. Part of that work is recognising when the first kind would do the job better.

A production agent is a stack of parts, and each one is a place where it can fail:

Part

The question it answers

A typical failure

Instructions

What is the agent for, and what must it never do?

Rules that live only in the prompt, with nothing enforcing them

Context

What does the model see when it decides?

Stale or conflicting information outranks the instruction

Model

Which model, and which version?

A provider update changes behaviour without a release

Tools

What can it change in other systems?

A tool granted for convenience does damage its job never needed

Permissions

Under whose authority does it act?

A shared service account with far more access than the task

Memory and retrieval

What does it remember and look up?

It retrieves a document the user isn't allowed to see

Workflow and state

Where is it in the task, and what happens next?

It carries on after a step failed

Human approval

Which actions wait for a person?

Approval becomes a rubber stamp

Evaluation

How good is it, measured how?

The test set only contains the cases the demo handled

Logging and monitoring

What happened on each run?

Nobody can reconstruct why it did what it did

Fallback and escalation

What does it do when unsure?

It guesses confidently instead of handing over

The paths below work through these parts in order. Throughout, the test of a design is whether each part has a check you can actually run.

How this curriculum is different

Four principles run through every path. They are opinions, and the articles argue for them rather than assume them.

  • Agents are systems, not prompts. Most agent failures in production come from tools, permissions, state and missing checks, not from wording. AI Agents Fail in Production in the Same Five Ways shows the pattern.

  • Verification drives the design. Write the goal, the expected behaviour, the failure conditions and how each would be detected before choosing an architecture. The method is set out in Verification-First Agent Design.

  • Autonomy is earned with evidence. An action gets more independence only when its failures can be detected and reversed. The autonomy ladder in What Is AI Agent Management? is the model used across the site.

  • Prompts are software artefacts. They are versioned, tested and regression-checked like code, with changes reviewed and evaluated before release.

The eight paths

The paths are ordered as a journey, but each stands alone. Articles marked coming are in production. The rest are live now.

1. Agent foundations

What an agent is, how the loop works, and how to decide whether you need one at all.

2. Prompt engineering as software engineering

Writing instructions a model can't misread, then testing and versioning them like code.

3. Context engineering

An agent's prompt is mostly assembled at run time. This path covers what goes into the context window, in what order, and how to stop it going stale.

4. Build an agent from scratch

One agent, built stage by stage, from defining the problem to monitoring it in production.

5. Agent architecture

The patterns (router, planner–executor, generator–critic, supervisor–worker) and when a single agent is enough.

6. Agent testing

How to test a system whose output changes every run, by checking what it did in the world rather than what it said.

7. Evaluation and operations

Golden datasets, judges, regression gates, telemetry, cost and drift: keeping an agent good after launch.

8. Governance and human oversight

Ownership, inventory, autonomy, change control, audit and the kill switch. The buyer's side of the same system.

Three ways in, depending on your job

If you test or assure software, start at path 6 with A Test Strategy for LLM Features, then read Verification-First Agent Design. Your existing skills transfer more directly than vendor material suggests. The new part is testing outcomes across repeated runs rather than a single expected output.

If you analyse requirements or run delivery, start at path 4 with Discovery for AI Features. Most agent failures trace back to a rule the people doing the job knew and nobody wrote down. Finding those rules is requirements work, and it is the part builders most often skip.

If you are accountable for the release decision, start at path 8 with What Is AI Agent Management? and the inventory article, then ask your teams the readiness questions below.

Is your team ready to run an agent?

Before an agent writes to a system of record, contacts a customer or touches money, your team should be able to answer yes to each of these with evidence, not intent:

  1. Is the agent on an inventory, with a named owner?

  2. Is its worst failure written down, and was that list reviewed by someone who does the job today?

  3. Does each of its powerful actions have a check that detects that failure, and does the check look at the system rather than the agent's report?

  4. Is the autonomy of each action recorded, with the evidence that justified it?

  5. Can you reconstruct any run from last week: what it saw, what it called and what changed?

  6. Has someone rehearsed stopping it, and reversing what it did?

Each "no" points to a path. Questions 2 and 3 are path 6 and Verification-First Agent Design. Questions 1, 4 and 6 are path 8. Question 5 is path 7.

Start here

Pick the agent your organisation is closest to putting into production and answer question 3 for its three most powerful actions. If any action has no check behind it, read Verification-First Agent Design next and run its 90-minute workshop before the go-live date, not after.

Sources

  • Anthropic (Schluntz, E. and Zhang, B.), Building effective agents, 19 December 2024. The distinction between workflows and agents.

Frequently Asked Questions

What is AI agent engineering?

The discipline of designing, building, testing and operating systems in which a language model chooses actions toward a goal, observes the results and continues. It covers the model's instructions and context, its tools and permissions, its memory and retrieval, and the testing, monitoring and human oversight that make its behaviour trustworthy.

What is the difference between an AI agent and a workflow?

In a workflow, code fixes the sequence of steps and the model fills in parts of it, such as classifying or drafting. In an agent, the model decides which steps to take and which tools to call. Workflows are easier to verify, so they are the better choice whenever the path is known in advance.

Where should I start learning to build AI agents?

If agents are new to you, start with the foundations path and the verification-first design method. If you already run an agent, start with the testing path and check that each of its powerful actions has a check behind it.

bottom of page