The symptom is familiar: someone built a useful agent, wired it up with an API key that had broad scope because that was what was available, and it now sits in production able to do considerably more than its job requires.
This is not carelessness. It is the default outcome of how agents get built, because the fast path to a working prototype is to use credentials you already have.
Why it happens by default
Credentials are inherited, not designed. The developer's own token, or a service account created years ago for an integration, becomes the agent's identity. Nobody scoped it, because scoping would have slowed the prototype down.
Tool lists grow and never shrink. A tool added during development to debug something stays in the list. Each one widens the blast radius.
There is no agent inventory. Agents get built inside teams, on platform tools, in notebooks, and in SaaS products that quietly shipped an agent feature. Nobody holds the list.
Conventional IAM does not model delegation. Your identity system knows about users and service accounts. It has no concept of "this agent, acting for this person, for this purpose, within these limits, until this date."
Why it matters more than ordinary over-permissioning
An over-permissioned human is bounded by working hours, attention, and intent. An agent is not. Three specific amplifiers:
Speed. An agent can execute thousands of actions in the time a human takes to make one mistake.
Prompt injection turns permissions into an attack surface. If an agent reads untrusted content — a web page, an email, an uploaded document, output from another agent — attacker-controlled text can direct it to use its tools. Every permission the agent holds is a permission the attacker effectively holds.
Nothing looks anomalous. The agent uses valid credentials to call endpoints it is authorized to call. Conventional logging shows an authorized service account doing authorized things.
The severity of any agent incident is set by the tool list, not by the sophistication of the attack. Reducing the tool surface is the single highest-return control available.
Step 1 — Inventory
You cannot scope what you have not listed. For each agent, record: what it does, what credential it runs as, which tools and endpoints it can reach, whether it can write or only read, who owns it, whether it processes untrusted input, and what the worst plausible outcome of a bad call is.
Expect surprises. Most organizations doing this for the first time find agents nobody in security knew about, several sharing a single powerful service account, and at least one with production write access built for a demo.
Step 2 — Scope down
For each agent, ask what it actually needs, not what it currently has.
- Split read and write. Most agents need to read far more than they write. Separate credentials, and require a higher bar for the write path.
- Cut the tool list. If a tool has not been called in ninety days, remove it. Agents also get measurably more reliable with fewer tools — this improves quality as well as security.
- Scope by resource, not by system. "Read tickets in the support queue" rather than "read the ticketing system."
- Add limits that match the job: transaction caps, rate limits, time-of-day restrictions, and an expiry date. An agent built for a Q3 project should not still hold credentials in Q2 of the following year.
- Give each agent its own identity. Shared service accounts make attribution impossible and mean revoking one agent breaks others.
Step 3 — Enforce at the action, not the login
Scoping is necessary and insufficient, because a correctly-scoped agent can still be manipulated into misusing the permissions it legitimately holds.
The durable control is an enforcement point between agents and the systems they touch, evaluating each consequential action against identity, delegated authority, policy, and behavioural baseline — then allowing, challenging, escalating to a human, blocking, or containing.
That is what Guardian Agent Defense does, and the architectural reason it sits in the request path rather than in the agent is that an agent cannot be trusted to police itself once its context has been poisoned.
At minimum, even without a dedicated control layer:
- Require human approval on irreversible actions — money movement, external communication, deletion, permission changes.
- Validate outputs against a schema and business rules before acting.
- Log the full chain: human, agent, permission, action, system, outcome.
- Alert on agents acting outside their normal pattern.
To work through this systematically against one deployment, use the 24-control AI agent security checklist — ungated, and derived from our Controlled Autonomy Standard.
Frequently asked questions
How do I find all the AI agents in my organization?
Combine API key and service-account audits — looking for keys calling model provider endpoints — with SaaS discovery, a review of AI features in already-approved vendor products, and an amnesty survey asking teams what they have built. The survey usually surfaces more than the tooling, provided people believe there is no penalty for answering honestly.
Can we just use our existing IAM for agents?
Partly. Existing IAM handles authentication and coarse authorization well, and you should absolutely use it. What it does not model is delegated authority — an agent acting on behalf of a specific person, for a specific purpose, with limits — or per-action policy evaluation. Those need an additional layer.
What is the minimum viable control if we have no budget?
Three things, in order: give every agent its own identity rather than a shared account, remove write access from any agent that does not demonstrably need it, and require human approval for irreversible actions. Those alone eliminate most realistic worst cases.
Should agents have standing credentials at all?
Ideally not. Short-lived, narrowly-scoped, per-task credentials are considerably safer than long-lived keys, and they make revocation trivial. This is more work to implement and is the right target state for any agent with write access.
How do we handle third-party agents from customers or vendors?
Treat them as untrusted by default and require them to identify themselves. The emerging pattern is a verified agent registry — known agents with declared purpose and permissions get defined access, unknown automated traffic gets limited access, and bulk extraction gets blocked. That is far more workable than trying to detect and ban all automation.
