top of page

Authentication: The Boring Problem That Quietly Kills AI Agents

Shawn West
Jul 8
8 min read

Updated: 37 minutes ago

Rewritten, 4 October 2026. An earlier version of this post led with a figure we could not trace to a reliable source. This version keeps what the evidence supports.

The scheduling agent worked for exactly seven days. It read a team's Google calendars, proposed meeting slots and booked rooms. Then every call to Google started failing at once. The team spent a morning on the model, the prompts and a recent dependency bump before someone checked the error text: invalid_grant.


The cause was one setting. The agent's OAuth app was still in "Testing" status in its Google Cloud project, and Google issues refresh tokens that expire in seven days to apps configured that way (Google Identity, OAuth 2.0 documentation). The pilot ran for six days. Production lasted seven.


(Composite scenario. The seven-day expiry is documented Google behaviour; the team and agent are illustrative.)


Nothing about that failure is clever, and that's why it catches teams. Authentication failures in agents get misdiagnosed as model problems because the model is the interesting part. This piece is for the engineer or QA lead deciding whether an agent's integrations are ready for unattended operation. It covers why agents break auth differently, the failure modes, and a check that would have caught the seven-day cliff before launch.


Why auth that works for people breaks for agents


Authentication systems were designed around a person. When a session expires, the person is redirected to log in. When a sensitive action triggers a second factor, they pick up their phone. When an integration disconnects, they reconnect it from a settings page.


An unattended agent can do none of that. It gets a 401 or an invalid_grant and has three options: recover programmatically, escalate to a human, or fail. If the tool wrapper doesn't make the failure obvious, there's a fourth, worse option. The agent treats the error as content and carries on as if the call had worked.


Every integration also brings its own credential model, with its own lifetime and its own ways to fail:


Service (as documented)

Credential behaviour

What breaks if you don't plan for it

GitHub App

Installation access tokens expire after 1 hour

Long-running tasks fail mid-way unless the token is regenerated

Slack (token rotation on)

Access tokens expire every 12 hours; the refresh token is replaced on each refresh; rotation can't be turned off once enabled

An agent that caches a token, or reuses an old refresh token, loses access

Google OAuth

Refresh tokens stop working if access is revoked, if unused for six months, on some password changes, if a per-account token limit is exceeded, or after 7 days for apps in Testing status

A pilot that lasts less than a week hides the cliff



Three services, three lifetimes, three refresh behaviours. None of that is exotic. What's new is that an agent may hold a dozen of these at once, and no person is watching any of them.


Illustrative calculation: if each of seven integrations is authenticated correctly 98% of the time when the agent needs it, the chance that all seven are healthy at once is about 87% (0.98⁷). Those aren't measured rates. The calculation shows why "each integration is mostly fine" doesn't add up to "the agent is fine".


Run this: list every service your agent authenticates to. For each one, write down the credential type, its documented lifetime, and what happens when it expires. Where you can't fill a cell, you've found a production incident that hasn't happened yet.


The failure modes, and how each one hides


These patterns aren't unique to AI. Agents make them more frequent and harder to see.


Expiry mid-task. A multi-step task starts with a valid token, and the token expires partway through. Depending on the wrapper, the task fails without a clear reason, recovers slowly, or produces output that looks like success.


Refresh that doesn't propagate. The refresh succeeds, but the agent keeps using a cached token, or the refresh happened in a different process. Every call uses stale credentials while the refresh logs look healthy.


Rotation that invalidates the old refresh token. With rotating refresh tokens, as in Slack's model, each refresh returns a new refresh token and the old one stops working. Two agent processes refreshing the same credential will race, and the loser is locked out.


Scope changes. The token is valid, but it no longer grants access to an endpoint the agent needs. A 403 handled like a 401 sends the agent into refresh loops that can't succeed.


Throttling that looks like auth. Some services start rejecting requests under load. An agent that retries aggressively can turn a throttle into what looks like an auth outage, and its refresh attempts are then throttled too.


Second factors and step-up auth. A sensitive operation demands re-verification that software can't provide. The error text reads like an instruction, and the agent may try another route rather than stopping.


Shared-credential lockouts. Several agents share one account. One agent's retries trigger a lockout, and every agent on that account fails at once, across unrelated workflows.


Run this: in a test environment, revoke one integration's credential while the agent is mid-task, and watch what it does. The pass condition is a clear, attributed escalation, something like "authentication failed for calendar, agent scheduling-01, human action required". If the agent reports success, retries without limit or invents an answer, you've found the gap.


The discovery move: ask about credential lifetime at intake


The seven-day cliff wasn't a coding mistake. It was a question nobody asked when the agent was scoped: what is the lifetime of every credential this agent depends on, and who or what renews it?


That question belongs on the agent's intake form, next to "what can it change" and "who owns it". Answering it forces someone to read each provider's token documentation before the pilot, not after the first outage. It also exposes the configuration traps that pilots hide, like Testing-mode expiry, rotation that can't be undone, and tokens that die after six months of disuse.


Run this: for every agent going into production, require a filled credential-lifetime table. Then check whether the pilot ran longer than the shortest refresh-token lifetime in it. If it didn't, the pilot never tested renewal.


Three layers of defence


Auth problems are solved in the layer between the agent and the service, not in the prompt. The model should never be trying to work out what an auth error means.


1. Proactive renewal. Renew tokens on a schedule ahead of expiry, not in response to a failure. Lifetimes are documented and predictable. Most OAuth client libraries support this; many agent frameworks don't turn it on by default. Where refresh tokens rotate, make a single component responsible for refreshing each credential, so concurrent processes don't race.


2. Health checks before work starts. Before a multi-step task begins, confirm that every integration it needs is healthy: the token is valid, the last call succeeded recently, and there's no spike in 401 or 403 responses. If something isn't healthy, defer or escalate before starting. That's cheaper than failing halfway and unpicking partial work.


3. Loud, structured failure. When recovery isn't automatic, emit a structured event naming the service, the scope, the agent and the provider's error, and route it to a person. Don't retry indefinitely, and never hand the model a broken call with an instruction to do its best.


Run this: count how many of the three layers your agent has today. Most stacks have only the third, and usually only after an incident.


Credential hygiene that pays off later


One credential per agent per service. Attribution, revocation and audit all depend on it. It's also what keeps one agent's retries from locking out every other agent.


Short-lived tokens by default. A leaked short-lived token is useful for hours, not until someone notices.


Refresh tokens stored separately. Keep them in a secrets manager, not in the agent's memory, prompts or logs. The current OAuth security best practice requires refresh tokens issued to public clients to be either sender-constrained or rotated (RFC 9700, January 2025).


Narrow scopes from the start. Broad scopes "for development" tend to survive into production, because narrowing them later breaks things nobody remembers depending on.


Logging at the credential layer. Log every credential use with agent, service, scope, time and result. When something goes wrong, this is the record you reconstruct from.


What to monitor


  • Auth success rate per agent and service pair, not aggregated. A dip for one agent against one service is the early warning that aggregate dashboards hide.

  • Refresh failures, tracked separately from call failures. A failing refresh means calls will start failing soon, even if they're fine now.

  • Time since the last successful call, per integration. Quiet can mean the agent didn't need the service, or that it gave up. Make the difference visible.

  • Scope and permission errors, routed separately. They're usually configuration, not outages, and belong with whoever owns access.

  • Days until each refresh token's known expiry, for providers with a fixed horizon.


Tradeoffs and exceptions


Per-agent credentials cost setup and admin time. For a single, low-risk, read-only internal assistant, a shared read-only credential may be an acceptable shortcut. Record it as a known exception, and revisit it before the agent gets write access.


Proactive renewal adds moving parts. A scheduler and a token store are more infrastructure to run. For short tasks against providers with long-lived tokens, refresh-on-failure with clear escalation may be enough. The test is whether a task can outlive its token.


Health checks add latency. For interactive agents, check asynchronously or cache recent health, rather than blocking every request.


Implementation reality


Make the credential-lifetime table a required section of agent intake, and its owner the agent's named owner. Add the mid-task revocation test to the release checklist for any agent with write access. Then put days-to-expiry for fixed-horizon tokens on the same dashboard as task success, so renewal problems show up before the cliff.


Credentials are where an agent's authority becomes real, which is why they belong in the controls stage of AI agent management, next to the autonomy level they are meant to enforce.


Authentication is one of the controls the AI Readiness Assessment probes under guardrails: whether there are limits on what the model can do without a human approving it, and whether untrusted input can make it act. Five minutes, nothing stored.


What to do next


Pick your longest-running agent and fill in the credential-lifetime table for it today. Check the shortest lifetime against how long the agent has actually run without a manual reconnect. If the agent hasn't yet outlived its shortest-lived credential, its renewal path has never been tested.


Final takeaway


Agents don't fail at authentication because auth is hard for AI. They fail because auth was built for people, credentials expire on schedules nobody wrote down, and agents remove the person who used to notice. Ask the lifetime question at intake, renew before expiry and fail loudly. Then the boring problem stays boring.


Sources


  • Google Identity, Using OAuth 2.0 to Access Google APIs, refresh token expiration (including the 7-day expiry for apps in Testing status). https://developers.google.com/identity/protocols/oauth2

  • GitHub Docs, Generating an installation access token for a GitHub App (tokens expire after 1 hour). https://docs.github.com/en/apps/creating-github-apps/authenticating-with-a-github-app/generating-an-installation-access-token-for-a-github-app

  • Slack, Using token rotation (access tokens expire every 12 hours; rotation can't be turned off once enabled). https://docs.slack.dev/authentication/using-token-rotation

  • IETF, RFC 9700, Best Current Practice for OAuth 2.0 Security, January 2025. https://www.rfc-editor.org/rfc/rfc9700



Related reading


bottom of page