Agent Governance Starts With an Inventory Nobody Has
The first hour of an agent incident is spent finding out what the agent is. That hour is the governance failure, and a policy document will not get it back.
At 9:40 on a Tuesday, the claims operations lead at a mid-sized insurer noticed that customer records were changing. Addresses were being merged, and duplicate policyholder entries were collapsing into single records. Some of those merges were wrong: two different people with the same name and postcode were now one customer, with one claims history.
The cause turned out to be a "data hygiene" agent. Someone in operations had built it on the CRM vendor's agent builder three months earlier. It was meant to flag likely duplicates for review. A configuration change the previous week had flipped it from suggest to apply.
The incident bridge asked the three questions every agent incident asks:
Who owns this? The builder had moved teams. The CRM admin team said they owned the platform, not what people built on it.
What else can it touch? Nobody knew. It ran under an integration service account with write access to the CRM, the document store and the outbound email connector.
How do we stop it without breaking something else? Disabling the service account would stop the agent. It would also stop the nightly claims export and two other integrations that used the same account. Nobody could say which.
It took 70 minutes to stop the agent safely, and nine days to unpick the wrong merges by hand.
(Composite scenario. The details are drawn from patterns that recur across regulated organisations, and are not a single verified incident.)
Nothing in that story involves a clever model failure. The agent did what its configuration said. The failure was that the organisation had granted authority it could not see, and so could not withdraw.
Why the policy-first instinct is attractive
When leadership decides agents need governing, the natural first move is a policy: an acceptable-use statement, an approval workflow and a review board. That instinct is reasonable. A policy is visible and fast to write, and it shows the organisation has taken the risk seriously. It is also what most frameworks seem to ask for first.
The trouble is that a policy governs the agents that come through the front door. In most enterprises today, agents arrive through four doors, and only one has a doorman:
Door | How an agent arrives | Who usually knows |
Engineering-built | A team builds an agent on an LLM API and deploys it through the normal pipeline | Engineering, sometimes the architecture review |
Vendor-embedded | A SaaS platform you already pay for ships an agent feature, sometimes switched on by an admin, sometimes by default in an update | The platform admin, if anyone |
Builder-made | Business teams use a low-code or vendor agent builder and connect it with an existing integration account | The person who built it |
Developer tooling | Coding agents on engineers' machines, connected to repositories, ticketing and cloud consoles through MCP servers | The individual developer |
A review board catches the first door. The insurer's agent came through the third. The fourth door, coding agents wired to internal systems through the Model Context Protocol, is growing quickly and is rarely counted at all. The NSA's AI Security Center warned in May 2026 that "MCP's rapid proliferation has outpaced the development of its security model" (NSA AISC, 2026).
Before you write the policy, count your doors. For each of the four, name the person who could tell you, today, how many agents came through it last quarter. If a door has no name against it, your policy does not reach that door.
What existing inventories miss: they key on the wrong thing
Most organisations that take AI risk seriously already have some kind of register: an AI use case inventory, a model registry or a CMDB entry per application. The NIST AI Risk Management Framework asks for exactly this: "Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities" (GOVERN 1.6, NIST AI 100-1). US federal agencies must publish AI use case inventories every year under OMB memo M-25-21 (OMB, 2025).
Those registers are necessary, and they share one design assumption that breaks for agents. They describe what a system is, not what it can do. A use case entry records the intent ("flag duplicate customer records"), the model and the business owner. It does not record that the system holds a credential that can write to three systems and send email.
For a conventional application, that gap is tolerable: what the code can do changes only when the code changes, and that goes through release control. For an agent, authority sits in places that change without a release:
Tool grants. Adding an MCP server or a connector gives the agent new verbs, with no code change.
Credentials. Agents often borrow an existing service account or a human's OAuth token, so they inherit everything that identity can do.
Autonomy settings. As at the insurer, one toggle moves an agent from suggest to apply.
OWASP's Top 10 for LLM Applications names the result Excessive Agency and gives it three root causes: excessive functionality, excessive permissions and excessive autonomy (OWASP LLM06:2025). Notice that none of the three is about the model. All three live in configuration that a use case register never sees.
The discovery move: build the agent inventory from the authority side, not the intent side. Don't start by asking teams "what agents do you have?" Start from the credentials and tool grants that exist, and trace each one back to whatever is using it.
To test your current register, take its three most recent AI entries. For each, can the register tell you which credential the system runs under? If it can't, it records intentions, not authority.
How agent authority actually accumulates
Authority accumulates in a predictable chain, and it is the same chain that produced the insurer's 70-minute stop:
Borrowed identity → inherited permissions → added tools → raised autonomy → unknown blast radius
Borrowed identity. The builder needs the agent to read the CRM. The quickest route is the integration account that already works. The agent now has every permission that account has, not just the ones the task needs.
Inherited permissions. That account was created years ago for a sync job and given broad write access "to be safe." Nobody revisits it, because it isn't new.
Added tools. Someone wants the agent to email customers when it merges records. The email connector is added. The agent's verbs now include send.
Raised autonomy. Review volume is high, so someone flips the agent from suggest to apply. That step is the one that turns the latent risk into an incident.
Unknown blast radius. Because the identity is shared, nobody can withdraw the agent's authority without also withdrawing other systems' authority. The kill switch exists, but it takes out the claims export too.
Every link in that chain is a reasonable local decision. No single step would fail a review, and that is why policy-first governance misses it. A review board sees a request at one point in time. Authority builds up between reviews.

A worked inventory entry: the same agent, before and after
Here is the insurer's agent as its use case register described it, and as an authority-keyed inventory card would describe it. (Composite, continuing the scenario above.)
Before — use case register entry:
CRM duplicate detection. Owner: Claims Operations. Model: vendor-hosted. Purpose: flag likely duplicate policyholder records for human review. Risk: low (advisory only).
Every word of that entry was accurate the day it was written. None of it was accurate the day of the incident.
After — authority-keyed inventory card:
Field | Value |
Agent ID | AGT-0147 · crm-dedupe |
Entry door | Builder-made (CRM vendor agent builder) |
Accountable owner | Head of Claims Operations (named person, not a team) |
Technical owner | CRM platform team, named engineer |
Identity | Shared integration account svc-crm-sync, also used by the claims export and 2 other integrations |
Tools / connectors | CRM records (read/write), document store (read/write), outbound email (send) |
Data touched | Policyholder PII, claims history (restricted) |
Unattended actions | Merge records, update addresses, send customer email |
Approval gate | Configurable; currently off (apply mode) |
Reversibility | Merge: compensable only by manual unpick. Email: irreversible |
Authority tier | 3 (irreversible actions on restricted data) |
Kill switch | Disable agent in builder console. Not tested. Identity revocation unsafe (shared) |
Logs | Builder console only, 30-day retention |
Last reviewed | At creation, 3 months ago |
Read the second card the way an incident commander would. Four fields would have changed what happened, and they are in bold: the shared identity, the gate being off, the irreversible actions, and the untested kill switch. A card like this, kept current, turns the 70-minute question into a 2-minute lookup. More importantly, it would have triggered a review the day the gate was switched off, because the agent's tier changed.
Ask yourself: if you wrote this card today for your most-used agent, how many fields would say unknown? That number is how far you are from being able to govern that agent.
How to evaluate your own position: the 15-minute test
You don't need a maturity model to know where you stand. Pick five agents at random, at least one from each door you can find, and time how long it takes to answer four questions for each:
Who is the accountable owner? A named person with authority to switch the agent off, not a team alias.
What can it change? Every system it can write to, and every external action it can take, such as sending, paying or publishing.
How do you stop it, and what else stops with it? The specific control, and whether using it affects anything else.
What did it do in the last 24 hours? Evidence from logs, not a description of intended behaviour.
Recommended threshold (ShiftQuality's recommendation, not an external standard): if you can't answer all four within 15 minutes for an agent, treat it as uninventoried, however many documents mention it. The number is deliberately tight. It forces the answers to come from a record rather than from asking around, which is the only way they will be available during an incident.
Score the result simply: the number of the five agents that pass. Two or fewer means you have a discovery problem, not a policy problem. Start with the sweep below.
The method: a four-week authority sweep
This is designed for a director to sponsor and a small working group to run: a security or identity engineer, a platform owner for each major SaaS system, and a quality or delivery lead to own the register. It produces a live inventory and a short list of agents that need action now.
Week 1 — Sweep the credentials, not the teams. Pull the evidence of authority from the systems that hold it:
Identity provider: export OAuth consent grants and app registrations, and filter for write, modify or send scopes and offline access. Also list service accounts created or used in the last 12 months.
Secrets manager: list API keys for LLM providers and for systems agents commonly call.
Repositories: search for committed MCP configuration. Project-level files have conventional names, such as .mcp.json, .cursor/mcp.json and .vscode/mcp.json, and most clients declare servers under an mcpServers key. In each repository, git ls-files | grep -E '(^|/)(\.mcp|\.cursor/mcp|\.vscode/mcp)\.json$' lists the files, and git grep -l '"mcpServers"' catches the rest. User-level configurations on laptops need an endpoint-management query instead.
SaaS admin consoles: for each major platform (CRM, ITSM, collaboration suite, HR system), list the agent features enabled and the custom agents built on it.
The output is a list of authority holders, not agents. Many will turn out to be ordinary integrations. That's fine, because the point is to start from what can act.
Week 2 — Trace and card. For each authority holder that is, or is used by, an agent, find the builder and fill in the inventory card above. Any field that nobody can fill is recorded as unknown, not left blank. Unknowns are the finding.
Week 3 — Tier and drill. Assign each agent an authority tier:
Tier | Authority | Minimum controls |
0 | Read-only, internal data | Card, named owner |
1 | Writes that are easily reversed (drafts, tags, internal tickets) | + dedicated identity, logs kept |
2 | Externally visible or compensable actions (customer email, record updates, refunds under a limit) | + approval gate or rate limit, kill switch tested |
3 | Irreversible actions, restricted data or regulated decisions | + human approval per action or batch, quarterly kill-switch drill, log retention set by legal requirement |
For every Tier 2 and Tier 3 agent, run a kill-switch drill: actually stop it in a controlled window, time it, and record what else broke. A kill switch that has never been pulled is an assumption, not a control.
Week 4 — Close the doors. Turn the sweep into a standing gate, so the inventory doesn't decay:
No new Tier 1+ agent gets a credential without a card. Enforce this where credentials are issued, in the identity provider and the secrets manager, not in a policy PDF.
Any change to an agent's tools, identity or autonomy setting re-triggers tiering. The toggle that caused the insurer's incident should create a ticket.
SaaS renewals and major updates include an "agent features enabled?" check by the platform owner.
You're done when every Tier 2–3 agent has a complete card, a dedicated identity and a drilled kill switch, and the credential gate is live.
Tradeoffs and where this changes
Option | Gains | Costs | Best suited for | Warning signs |
Central register, owned by one team | One source of truth, consistent tiering | Becomes a bottleneck; builders route around it | Regulated orgs with few agent builders | Card backlog grows; shadow agents appear on shared accounts |
Federated cards, central standard | Scales with builder count; owners keep cards current | Uneven quality between teams | Large orgs with many platforms | Cards with stale "last reviewed" dates |
Inventory by class, not instance | Practical for thousands of near-identical agents | Loses per-instance blast radius | Coding agents on developer machines; personal assistants | A class card covering agents that have been given extra tools |
Exceptions worth making explicitly:
Personal, read-only assistants (Tier 0) don't need individual cards. Inventory the class and its allowed data sources. Carding every one creates friction that pushes people to unsanctioned tools.
Developer coding agents are best handled by class plus configuration policy: approved MCP servers, no production credentials on laptops, and repository rules that flag committed MCP configs for review. Per-laptop cards are not realistic.
Vendor-embedded agents you cannot configure in depth still get a card. The card is where you record that you can't see inside them. That is also the evidence for the conversation with the vendor, and see why your AI vendor's compliance doesn't make you compliant.
The tradeoff underneath all three is friction against coverage. A register that is painful to use gets bypassed, and an agent on a borrowed identity is the most expensive kind to discover later.
Implementation reality: where this breaks
The inventory goes stale. A spreadsheet updated by hand is accurate on the day of the sweep and wrong a month later. Prevent this by generating as much of the card as possible from the systems in Week 1. Identity, scopes and connectors can be pulled automatically. Owner and tier need a human, so review those fields on a schedule.
Owners leave. Cards that name a team, or a person who moved, are the most common silent failure. Make "accountable owner left" an automatic trigger. Joiner-mover-leaver processes already exist for people, so extend them to the agents those people own.
Shared identities resist fixing. Splitting svc-crm-sync into four identities means touching four integrations, and nobody budgets for that. Do it first for Tier 3 agents. The shared identity is what turns a 2-minute stop into a 70-minute one.
Logs exist but nobody can reach them. Builder consoles often keep short logs that only platform admins can see. NSA's MCP guidance calls for logging each tool request, who made it and what resulted (NSA AISC, 2026). If you deploy high-risk systems in the EU, Article 26 of the AI Act requires deployers to keep automatically generated logs for at least six months and to assign human oversight to people with the competence and authority to exercise it (EU AI Act, Art. 26). Those high-risk obligations now apply from 2 December 2027 for Annex III systems, after the July 2026 Digital Omnibus (Orrick, 2026). That is the time you have to build an inventory you can show a regulator, not a reason to wait. For the wider logging question, see production observability for agents.
The business case gets questioned. Gartner predicted in June 2025 that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls (Gartner, 2025). An inventory is also how you find out which of your agents are earning their authority and which can be retired, and decommissioning is a governance outcome too (NIST GOVERN 1.7).
The inventory is step zero of a larger discipline. What is AI agent management? sets out the full lifecycle it feeds: the workflow map, the digital-worker definition, the autonomy level and the controls each agent needs.
Once you have the inventory, the AI Readiness Assessment scores an agent across evaluation, guardrails, observability and governance and shows its thinnest area. Run it once per agent, starting with the one that can write to the most systems.
What to do next
This week, run the 15-minute test on five agents. Use whatever you can find, and include at least one vendor-embedded agent and one that uses MCP. If you need a primer on how MCP connects agents to tools before you search for it, start with MCP: the protocol that connects AI to everything.
Write down the number that passed. Then take the result to whoever owns your identity provider and ask for one export: every OAuth grant and service account with write scope that has been used in the last 90 days. That export is Week 1 of the sweep, and it's the first artefact in your agent inventory.
Final takeaway
Agent governance is usually framed as a policy problem. It's a visibility problem first. An agent's risk lives in its credentials, tool grants and autonomy settings, which change without a release and are invisible to the registers most organisations already keep. Build the inventory from the authority side, tier each agent by what it can do unattended, and pull every kill switch once before you need it. A policy is enforceable only for the agents you can list.
Sources
NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023 — GOVERN 1.6 (inventory) and GOVERN 1.7 (decommissioning). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
NIST AI RMF Playbook, GOVERN function. https://airc.nist.gov/airmf-resources/playbook/govern/
OWASP GenAI Security Project, LLM06:2025 Excessive Agency. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
NSA Artificial Intelligence Security Center, Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation, Cybersecurity Information Sheet U/OO/6030316-26, May 2026. https://www.nsa.gov/Portals/75/documents/Cybersecurity/CSI_MCP_SECURITY.pdf
Office of Management and Budget, M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, April 2025. https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-21-Accelerating-Federal-Use-of-AI-through-Innovation-Governance-and-Public-Trust.pdf
Regulation (EU) 2024/1689 (AI Act), Article 26, Obligations of deployers of high-risk AI systems. https://artificialintelligenceact.eu/article/26/
Orrick, EU AI Act Update: Digital Omnibus Finalizes 8 Compliance Changes, July 2026 (Regulation (EU) 2026/1744). https://www.orrick.com/en/Insights/2026/07/EU-AI-Act-Update-Digital-Omnibus-Finalizes-8-Compliance-Changes
Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027


