AI Agent Engineering: Build, Test and Govern Agents That Hold Up in Production
An agent is not a prompt with ambitions. It is a system that takes actions in other systems, and engineering one means being able to say, for every action it takes, how you would know it was wrong.
Consider a situation many quality leads are now in. An agent has been built to handle part of an operational process: reading requests, updating records, drafting replies. The demo went well and the sponsor wants it live next month. The quality lead is asked to "test it", and finds there is no specification to test against. There is a prompt, a list of tools and a handful of examples that worked. What the agent must never do, which of its actions can be undone, and what evidence would show it is behaving: none of that was written down, because the design started from what the agent could do.
That is the gap this curriculum is built around. There is no shortage of material on how to wire a model to some tools. There is far less on how to decide what an agent should be allowed to do, how to prove it is doing it, and how to keep proving it after the model, the prompt or the permissions change. ShiftQuality's agent engineering paths cover both halves, and they put the second half first.
What agent engineering covers
Anthropic's widely cited guide draws a distinction worth adopting. Workflows are systems where language models and tools are orchestrated through predefined code paths. Agents are systems where the model dynamically directs its own process and tool use (Anthropic, 2024). Agent engineering is the work of building the second kind safely. Part of that work is recognising when the first kind would do the job better.
A production agent is a stack of parts, and each one is a place where it can fail:
Part | The question it answers | A typical failure |
Instructions | What is the agent for, and what must it never do? | Rules that live only in the prompt, with nothing enforcing them |
Context | What does the model see when it decides? | Stale or conflicting information outranks the instruction |
Model | Which model, and which version? | A provider update changes behaviour without a release |
Tools | What can it change in other systems? | A tool granted for convenience does damage its job never needed |
Permissions | Under whose authority does it act? | A shared service account with far more access than the task |
Memory and retrieval | What does it remember and look up? | It retrieves a document the user isn't allowed to see |
Workflow and state | Where is it in the task, and what happens next? | It carries on after a step failed |
Human approval | Which actions wait for a person? | Approval becomes a rubber stamp |
Evaluation | How good is it, measured how? | The test set only contains the cases the demo handled |
Logging and monitoring | What happened on each run? | Nobody can reconstruct why it did what it did |
Fallback and escalation | What does it do when unsure? | It guesses confidently instead of handing over |
The paths below work through these parts in order. Throughout, the test of a design is whether each part has a check you can actually run.
How this curriculum is different
Four principles run through every path. They are opinions, and the articles argue for them rather than assume them.
Agents are systems, not prompts. Most agent failures in production come from tools, permissions, state and missing checks, not from wording. AI Agents Fail in Production in the Same Five Ways shows the pattern.
Verification drives the design. Write the goal, the expected behaviour, the failure conditions and how each would be detected before choosing an architecture. The method is set out in Verification-First Agent Design.
Autonomy is earned with evidence. An action gets more independence only when its failures can be detected and reversed. The autonomy ladder in What Is AI Agent Management? is the model used across the site.
Prompts are software artefacts. They are versioned, tested and regression-checked like code, with changes reviewed and evaluated before release.
The eight paths
The paths are ordered as a journey, but each stands alone. Articles marked coming are in production. The rest are live now.
1. Agent foundations
What an agent is, how the loop works, and how to decide whether you need one at all.
AI Agents: What They Are and Why They Keep Failing: the working definition, and the failure pattern behind most incidents.
What Is an Agent (With Code): the observe, decide, act loop in a few dozen lines.
What LLMs Can't Do (And Why That Matters): the limits that should shape every design.
LLM Application Patterns: From Simple Completions to Reasoning Systems: where agents sit among the simpler designs.
Verification-First Agent Design: the method the rest of the curriculum uses.
Coming: agents vs workflows vs automation vs RPA, with a decision tree; the anatomy of a production agent.
2. Prompt engineering as software engineering
Writing instructions a model can't misread, then testing and versioning them like code.
Set Up Your Prompting Workspace and Write Your First System Prompt
Track Prompt Regression Over Time: the step most prompt guides leave out.
Coming: the Prompt Spec, a nine-part structure for prompts that ship with their own tests; prompt versioning and change control.
3. Context engineering
An agent's prompt is mostly assembled at run time. This path covers what goes into the context window, in what order, and how to stop it going stale.
When Prompting Isn't Enough: Fine-Tuning, RAG, and Knowing Your Options
Coming: a practitioner's guide to context engineering, covering budgets, instruction hierarchy, pollution and long-running agents.
4. Build an agent from scratch
One agent, built stage by stage, from defining the problem to monitoring it in production.
Discovery for AI Features: Scoping Before You Prompt: stage one, before any code.
Build a Single-Tool Agent, then Add Multiple Tools to an Agent
Function Calling and Tool Use and MCP: The Protocol That Connects AI to Everything
Agent Memory and State and RAG Architecture: When Your LLM Needs Real Data
Authentication: The Boring Problem That Quietly Kills AI Agents
Coming: a fifteen-stage series built on one running example, a requirements-intake triage agent, with its full blueprint.
5. Agent architecture
The patterns (router, planner–executor, generator–critic, supervisor–worker) and when a single agent is enough.
Build a Planning Agent: the planner–executor pattern in practice.
Coming: a field guide to agent patterns, each with its signature failure mode and how to test for it.
6. Agent testing
How to test a system whose output changes every run, by checking what it did in the world rather than what it said.
Coming: how to test an AI agent; testing non-determinism honestly; permission, boundary and escalation testing.
7. Evaluation and operations
Golden datasets, judges, regression gates, telemetry, cost and drift: keeping an agent good after launch.
Automated Grading with Rules and LLM-as-Judge for Subjective Criteria
Monitor for Prompt Drift and Continuous Evaluation in Production
Cost-Aware Evaluation and Running LLMs in Production: Cost, Latency, and the Tradeoffs Nobody Warns You About
Coming: evaluating AI agents end to end; an agent observability checklist.
8. Governance and human oversight
Ownership, inventory, autonomy, change control, audit and the kill switch. The buyer's side of the same system.
What Is AI Agent Management?: the eight-stage lifecycle and the autonomy ladder.
Coming: where human approval belongs; the agent governance framework, from register to retire.
Three ways in, depending on your job
If you test or assure software, start at path 6 with A Test Strategy for LLM Features, then read Verification-First Agent Design. Your existing skills transfer more directly than vendor material suggests. The new part is testing outcomes across repeated runs rather than a single expected output.
If you analyse requirements or run delivery, start at path 4 with Discovery for AI Features. Most agent failures trace back to a rule the people doing the job knew and nobody wrote down. Finding those rules is requirements work, and it is the part builders most often skip.
If you are accountable for the release decision, start at path 8 with What Is AI Agent Management? and the inventory article, then ask your teams the readiness questions below.
Is your team ready to run an agent?
Before an agent writes to a system of record, contacts a customer or touches money, your team should be able to answer yes to each of these with evidence, not intent:
Is the agent on an inventory, with a named owner?
Is its worst failure written down, and was that list reviewed by someone who does the job today?
Does each of its powerful actions have a check that detects that failure, and does the check look at the system rather than the agent's report?
Is the autonomy of each action recorded, with the evidence that justified it?
Can you reconstruct any run from last week: what it saw, what it called and what changed?
Has someone rehearsed stopping it, and reversing what it did?
Each "no" points to a path. Questions 2 and 3 are path 6 and Verification-First Agent Design. Questions 1, 4 and 6 are path 8. Question 5 is path 7.
Start here
Pick the agent your organisation is closest to putting into production and answer question 3 for its three most powerful actions. If any action has no check behind it, read Verification-First Agent Design next and run its 90-minute workshop before the go-live date, not after.
Sources
Anthropic (Schluntz, E. and Zhang, B.), Building effective agents, 19 December 2024. The distinction between workflows and agents.
Frequently Asked Questions
What is AI agent engineering?
The discipline of designing, building, testing and operating systems in which a language model chooses actions toward a goal, observes the results and continues. It covers the model's instructions and context, its tools and permissions, its memory and retrieval, and the testing, monitoring and human oversight that make its behaviour trustworthy.
What is the difference between an AI agent and a workflow?
In a workflow, code fixes the sequence of steps and the model fills in parts of it, such as classifying or drafting. In an agent, the model decides which steps to take and which tools to call. Workflows are easier to verify, so they are the better choice whenever the path is known in advance.
Where should I start learning to build AI agents?
If agents are new to you, start with the foundations path and the verification-first design method. If you already run an agent, start with the testing path and check that each of its powerful actions has a check behind it.


