top of page

Designing Tool Contracts and Least Privilege for AI Agents

Shawn West
12 minutes ago
13 min read

A tool is the point where an agent's decision becomes a change in someone else's system. Whatever the model must never do has to be impossible at that point, not just discouraged in the prompt.

The example in this article is a composite scenario built from patterns common to finance operations. It is not a single real organisation, and its figures are illustrative. The code is real: the test and replay output below were produced by running it.

At month-end close, the financial controller at a mid-sized distributor found two reversing journal entries for the same supplier invoice, each for 18,420.00. Both had been drafted by the company's duplicate-invoice investigator, an agent whose whole job is finding invoices booked twice. Two different clerks had approved them. The duplicate had been reversed twice, and the supplier balance was understated by exactly the amount the agent had been asked to correct.

Two weeks earlier, after three months as a read-only investigator and an evaluation that cleared the promotion bar (told in Evaluating AI Agents), the agent had been promoted to Level 2: it could draft the correcting entry, and a clerk would approve each draft before it posted.

The review found two things. The ERP had been slow that day: the request to create the draft committed, the response timed out, and the agent's loop, configured to retry a failed tool call once, called the tool again. Nothing in the tool could tell a retry from a new request. The second finding was worse. To ship quickly, the developer had wired the drafting tool to the ERP integration account the accounts-payable automation already used, and that account could post journal entries in all three legal entities. Between the agent and posting stood one sentence in its system prompt.

Why the prompt and the schema feel like enough

The team had followed the documentation. The tool had a detailed description, which Anthropic's tool-use docs call "by far the most important factor in tool performance" (Anthropic, n.d.-a). It had a JSON schema, and strict mode "guarantees that tool inputs will always match your schema exactly" (Anthropic, n.d.-b). And the system prompt said, clearly, never to post.

None of that addresses what went wrong. A description shapes which tool the model picks; a schema guarantees the call's shape. Neither says what happens when a correct call is made twice, or what else the credential behind the tool can reach. Choosing an action is the model's job. Deciding whether the action is allowed, safe to repeat and within limits is a different job, and nobody had been given it. OWASP's entry on excessive agency traces the risk to excessive functionality, permissions and autonomy, and recommends "complete mediation": "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not" (OWASP, 2025).

The test: for one agent, list every rule it must never break. Next to each, write where it is enforced. Any rule whose only enforcement point is the prompt is a request, not a control.

A schema checks shape; a contract checks consequence

When an agent calls a tool, three things travel together: the model's choice, the arguments, and the authority of the credential the tool uses. The prompt influences the first. The schema validates the second. The third is usually inherited from whatever account was convenient when the integration was built, and it decides the worst case.

Retries make the gap concrete. A timeout after a write leaves the caller unable to tell "it didn't happen" from "it happened and the reply was lost". Payments APIs handle this with idempotency keys. In Stripe's words, with a key "you can safely repeat the request without risk of creating a second object or performing the update twice", and the server "compares incoming parameters to those of the original request and errors if they're not the same" (Stripe, n.d.). Agents add retries of their own: Anthropic's docs note that after an invalid call, Claude "will retry 2-3 times with corrections" (Anthropic, n.d.-b), on top of whatever the loop does. The drafting tool had no key, so the retry was a new request.

Discovery would have caught both failures with two intake questions nobody asked: what happens if this tool is called twice with the same arguments? and what is the most powerful thing this tool's credential can do? The second was answerable from data the team already had. The ERP's role report for the integration account listed posting rights in every entity. Nobody pulled it, because the account was "already approved".

A tool contract answers both questions in writing and is enforced by the code that runs the tool. It has seven clauses.

Clause

What it says

The failure it prevents

1. Registered contract

No contract, no call

An unreviewed tool appearing in the agent's list

2. Scope

This tool needs exactly one scope, held by this agent's own credential

A shared account lending the agent powers it was never granted

3. Side-effect class

read, draft, write or irreversible, with the autonomy level each needs

A drafting agent that can, in practice, execute

4. Typed inputs

Known fields only, validated values, money as decimals

Schema-valid nonsense: wrong entity, float rounding, extra flags

5. Limits

Per-call amount cap and calls per hour

One bad decision repeated a thousand times

6. Idempotency

Every write carries a key that reaches the system of record

Retries that duplicate writes

7. Output and audit

Results checked against the output schema; every call logged, refusals included

Silent partial results; no trace of what was attempted

The Model Context Protocol specification asks for much of this: servers "MUST" validate all tool inputs, implement proper access controls, rate limit tool invocations and sanitize tool outputs; clients "SHOULD" log tool usage for audit purposes; and clients "MUST consider tool annotations to be untrusted unless they come from trusted servers" (MCP, 2025). A tool that describes itself as read-only hasn't proved it. A contract meets those requirements in a form you can test.

Side-effect class is the clause that connects tools to autonomy

On the autonomy ladder in What Is AI Agent Management?, Level 1 recommends, Level 2 drafts for a person to approve, Level 3 executes approved low-risk actions after deterministic checks, and Level 4 runs a bounded workflow. A promotion only means something if the tools enforce it, and classifying each tool by its side effect is how they do:

Side-effect class

Example in the finance agent

Who may call it

Read

Find an invoice; list ledger lines

Level 1 and up

Draft

Create a draft reversing entry

Level 2 and up; each draft approved by a person

Write (reversible)

Put a supplier invoice on hold

Level 3 and up, within limits, with a tested reversal

Irreversible

Post a journal entry; release a payment

Nobody unaided: a named human approver, checked at the tool

Two rules follow. The agent's credential carries its approved level and the gateway checks it on every call, so a demotion is one changed value that a test can confirm. And irreversible actions keep a human gate at the tool, not in the agent's workflow, so no prompt change or orchestration bug can route around it. OWASP says the same: "Utilise human-in-the-loop control to require a human to approve high-impact actions" (OWASP, 2025).

Human review alone wasn't enough, though. Level 2 review missed the duplicate because each clerk saw one plausible draft, and per-item review can't see duplicates across items. The idempotency key can.

The test: list your agent's tools and give each a side-effect class. Then compare the highest class with the level the agent was approved for. An agent approved at Level 2 that holds a write or irreversible tool is operating above its approval.

Least privilege means one agent, one identity, one scope per tool

NIST defines least privilege as the principle that "each entity is granted the minimum system resources and authorizations that the entity needs to perform its function" (NIST, n.d., citing SP 800-53 Rev. 5). For agents, the entity is the agent itself and the function is each tool, because an agent's function changes every time someone adds one. Five rules follow:

  • A dedicated identity per agent. Never a person's account, never a shared integration or admin account. If two agents share a credential, the audit log can't say which acted, and revoking one revokes both.

  • Scopes per action, not per system. The finance agent needed erp.invoice.read and erp.journal.draft. "ERP access" is not a scope.

  • Deny by default. A tool without a registered contract is refused, however it got into the agent's tool list.

  • Separate read from write. In OWASP's mailbox example, an assistant that only needed to read email held an extension that could also send it, and a crafted email turned that into exfiltration; the fix was a read-only extension with a read-only OAuth scope (OWASP, 2025). Anthropic suggests consolidating related operations into one tool with an action parameter (Anthropic, n.d.-a). If you do, each action still needs its own permission check.

  • Assume untrusted input. If the agent reads supplier emails or invoice PDFs, outsiders can put text in front of it. MCP Is Your New Attack Surface shows how that text reaches tools; least privilege decides how much damage it can do.

The test: export the real permission list for every credential your agent's tools use (the ERP role report, the cloud identity's policy). Mark each permission the tools actually call. Every unmarked write permission is a fix, and the first one is the most powerful.

The drafting tool, rebuilt as a contract

In the rebuild, a small gateway (about 190 lines of standard-library Python) sits between the agent and the ERP, and every tool is registered with a contract. This is the drafting tool's declaration and the investigator's credential:

gw.register(Contract("draft_reversal", "draft", "erp.journal.draft",
                     inputs={"entity": one_of(*ENTITIES), "invoice": ref("INV-"), "amount": money},
                     outputs=("draft_id", "status"),
                     max_amount=Decimal("25000.00"), max_calls_per_hour=30),
            lambda a, k: erp.create_draft(a["entity"], a["invoice"], a["amount"], external_ref=k))

# The investigator at Level 2: it can read and draft. It has no posting scope at all.
INVESTIGATOR = Credential("dup-invoice-investigator",
                          frozenset({"erp.invoice.read", "erp.journal.draft"}), level=2)

The gateway checks in this order. Each refusal returns a stable code, whether retrying can help, and what to do next, following the advice to say "what went wrong and what Claude should try next" (Anthropic, n.d.-b).

c = self.contracts.get(tool)
if c is None:                                               # deny by default
    raise NotPermitted(f"no contract registered for tool '{tool}'")
if c.scope not in cred.scopes:                              # least privilege, per tool
    raise NotPermitted(f"credential for {cred.agent_id} lacks scope {c.scope}")
floor = SIDE_EFFECTS[c.side_effect]
if floor is None:                                           # irreversible: a person, every time
    if not approval or approval.get("approved_by") in (None, cred.agent_id):
        raise ApprovalRequired(f"{tool} is irreversible and needs a named human approver")
elif cred.level < floor:
    raise NotPermitted(f"{tool} needs autonomy level {floor}; agent is at level {cred.level}")
clean = _validate(c, args)
if c.max_amount is not None and clean.get("amount", 0) > c.max_amount:
    raise LimitExceeded(f"amount {clean['amount']} exceeds the per-call limit of {c.max_amount}")
if c.needs_key:
    if not key:
        raise InvalidInput(f"{tool} writes; an idempotency_key is required")
    seen = self.keys.get(key)
    if seen:
        if seen[0] != _digest(clean):
            raise Conflict(f"key {key} was first used with different arguments")
        return dict(seen[1], replayed=True)                 # a retry, not a new write
self._rate_check(cred, c)
try:
    out = self.handlers[tool](clean, key)                   # the key travels to the system of record
except TimeoutError as e:
    raise Unavailable(f"{tool} timed out; outcome unknown ({e})")

Two details carry the weight. The idempotency key is derived from the case (dup-invoice-investigator:INV-88231:reverse:v1), not generated per call, so a retry carries the same key. And it is passed to the ERP as an external reference the ERP refuses to duplicate. A gateway cache alone can't cover a lost response, because the gateway never learned the result; the system of record has to recognise the key.

Replaying the incident against both designs, with the same lost response:

without a contract: 2 drafts for INV-88231 -> ['JE-D0001', 'JE-D0002']
with a contract:    1 draft  for INV-88231 -> ['JE-D0001']
audit:
  draft_reversal  key=dup-invoice-investigator:INV-88231:reverse:v1  outcome=unavailable
  draft_reversal  key=dup-invoice-investigator:INV-88231:reverse:v1  outcome=ok
agent tries to post: not_permitted (credential for dup-invoice-investigator lacks scope erp.journal.post)

The last line fixes the second finding. The investigator can't post because its credential lacks the scope, not because the prompt asked. Posting belongs to a separate service identity, and only with a named clerk's approval.

Every clause becomes a boundary test

How to Test an AI Agent treats boundary tests as their own layer: they check that prohibited actions are blocked by the tool layer, and they're deterministic because no model is involved. A contract gives you at least one per clause. Here are the fourteen for the finance agent's tools, with their actual output:

ok    test_agent_cannot_post_even_with_a_valid_draft
ok    test_amount_over_the_per_call_limit_escalates
ok    test_every_call_is_audited_including_refusals
ok    test_float_amount_rejected_with_a_fix
ok    test_level_1_agent_with_draft_scope_still_cannot_draft
ok    test_posting_needs_a_human_approver_who_is_not_the_agent
ok    test_read_only_credential_cannot_draft
ok    test_result_missing_a_contracted_field_is_treated_as_failure
ok    test_retry_after_lost_response_creates_one_draft
ok    test_same_key_with_different_arguments_is_a_conflict
ok    test_thirty_first_draft_in_an_hour_is_refused
ok    test_unknown_entity_and_extra_fields_rejected
ok    test_unregistered_tool_is_refused
ok    test_write_without_a_key_is_refused
----------------------------------------------------------------------
Ran 14 tests in 0.002s

OK

A passing suite proves little unless it fails when a control is removed. We ran four mutants, each deleting one control from a copy of the code. Removing the scope check failed two tests, including the posting test. Removing the level check, the key requirement or the ERP's external-reference dedupe each failed the test written for it. The dedupe mutant is the instructive one: with the gateway's key cache intact, the lost-response test still produced two drafts.

No test here asks the model to behave. Those belong in the trajectory and adversarial layers, run k times. Boundary tests run once per commit, because what they test is code.

How to audit your own agent's tools this week

Take the agent with the most write access and answer these with evidence:

  1. Which credential does each tool use, and is it used by anything other than this agent?

  2. What is the most powerful thing each credential can do, according to the system's own permission report?

  3. What happens if each write tool is called twice with the same arguments? Show the test.

  4. Which tools are irreversible, and where in code is the human approval checked?

  5. Does the agent's credential carry its approved autonomy level, and does a demotion test exist?

  6. Does the audit log record refused calls, and can you list last week's?

Fix a "don't know" on question 2 first. It is the gap that turned a drafting agent into a potential posting agent in the composite.

Tradeoffs and exceptions

Choice

Gains

Costs

Use when

Gateway in front of every tool

One place to enforce and test; uniform audit

Another component to own and run

Any agent with write tools

Credential per agent

Clean audit; independent revocation

Identity admin; secrets to rotate

Always, for agents that write

Idempotency at the system of record

Survives lost responses

Needs an external-reference field or dedupe API

Every draft and write tool

Human gate at the tool

Cannot be routed around

Approval queues; reviewer fatigue

Irreversible or high-value actions

Fine-grained scopes

Small blast radius

Some systems only offer coarse roles

Wherever the system supports it

Exceptions. A read-only agent over non-sensitive data needs less: its own identity, deny by default and input validation still apply; idempotency and approval gates are moot. Where a legacy system only offers an all-powerful service account, don't give it to the agent. Put a narrow API in front that exposes only the agent's operations, or keep the agent at Level 1 for that system. A coding agent in a disposable sandbox can hold broad permissions inside it, because the sandbox is the boundary; what leaves it, such as a push, goes through a contract.

How tool contracts go wrong

  • Contracts in a document, not in code. A permissions table nothing enforces is a role card nobody checks. The test is the contract.

  • Random idempotency keys. A fresh key per call makes every retry look new. Derive it from the case and the action.

  • Approval the agent can supply. If the approver field accepts the agent's own identity, the gate is decorative. Ours rejects self-approval, and a test proves it. It doesn't yet check the approver against the staff directory, and a production gate should.

  • Splitting to dodge limits. An agent capped at 25,000 per call may propose two drafts of 12,500. Ours only caps per call; a per-case cumulative limit, with its own test, closes the gap.

  • Logging only successes. Refused calls are the early warning that a prompt change or injected text is pushing the agent somewhere new.

Who owns what

The agent's owner signs the role card listing each tool and its side-effect class. The platform team owns the gateway and issues one identity per agent. Security reviews scopes against the system's real permission report at each promotion. QA owns the boundary tests and mutation check, run in CI on every change to a tool, contract or credential. A promotion is approved as a change to the credential's level, with the boundary tests as evidence. This is the design-time half of Verification-First Agent Design: each named failure gets a check, and for tools the check lives in the contract.

What to do next

Pick your agent's most consequential write tool. Write its contract as the seven clauses above, pull the real permission report for the credential behind it, and write three tests: a retry after a lost response produces one write, a credential without the scope is refused, and an irreversible call without a named approver is refused. Then delete the prompt sentence that forbids misuse and run the tests again. If they still pass, the control lives in the tool, which is where it belongs.

Final takeaway

An agent can only do what its tools allow and its credentials permit. Design each tool as a contract, give each agent an identity scoped to its role, and test every clause where the call meets the system. That gives you limits you can show an auditor, which a prompt can't.

Sources

  • OWASP GenAI Security Project, LLM06:2025 Excessive Agency, 2025. Three root causes; minimise extensions and permissions; complete mediation; human approval for high-impact actions; rate limiting; the read-only mailbox example.

  • Model Context Protocol, Specification 2025-11-25: Tools. Server MUSTs (validate inputs, access controls, rate limits, sanitize outputs); client SHOULDs (confirmation, audit logging); annotations untrusted; tool execution errors versus protocol errors; output schema conformance.

  • Anthropic (n.d.-a), Define tools, Claude API documentation, accessed 9 October 2026. Detailed descriptions as the most important factor in tool performance.

  • Anthropic (n.d.-b), Handle tool calls, Claude API documentation, accessed 9 October 2026. Errors with is_error; instructive error messages; strict tool use guarantees schema-conformant inputs.

  • Stripe, Idempotent requests, API reference, accessed 9 October 2026. Safe retries with idempotency keys; parameter comparison on reuse.

  • NIST Computer Security Resource Center, Glossary: least privilege, definition from NIST SP 800-53 Rev. 5.

Frequently Asked Questions

What is a tool contract for an AI agent?

A written, enforced agreement for each tool an agent can call: typed and validated inputs, an output schema, error types the agent can act on, an idempotency key for every write, a side-effect class (read, draft, write or irreversible), blast-radius and rate limits, and an audit record. The gateway that runs the tool enforces it, whatever the model asks for.

How do you apply least privilege to AI agents?

Give each agent its own identity and a credential scoped to the tools on its role card, never a shared admin or integration account. Deny any tool without a contract, separate read tools from write tools, and require a named human approver at the tool for irreversible actions.

Is a strict JSON schema enough to make agent tool calls safe?

No. Strict schemas guarantee the shape of a tool call, not its authority or consequence. A schema-valid call can still repeat a write, exceed a sensible amount or reach a system the agent should never touch. Those are enforced by scopes, limits, idempotency and approval at the tool.

Part of the Agent Engineering curriculum: eight paths from prompts and context to testing, evaluation and governance.

bottom of page