The blast radius of AI agents is what one of them could break, leak or spend if it went as wrong as possible. Credentials and reach set it, and the agent’s behavior does not. That makes it a quantity you can estimate before granting access: one row per grant, a worst case per row, and a time to stop.
I wrote this for the person who signs off on production access. Someone on your team wants credentials for an agent, and “can it be trusted” is a question nobody can answer. “What can it reach, what is the worst single action on each item, and how long until we stop it” has answers you can write down today. This post gives you the sheet, the arithmetic and a drill for the stop.
What is the blast radius of AI agents, and what sets it?
The blast radius of AI agents is set by what they hold: credentials, and whatever their process can reach. Chapter 17 of the book (in the full book) defines the blast radius as “everything a step, a run, or an agent could break or leak if it went as wrong as possible.” The section “Least Privilege, Sandboxing, and Human Approval” frames a deployment as “a bet with two components: how likely a failure is, and how much a failure costs.” Blast radius is the second component.
One vendor’s containment engineers describe the same split and say which half moves which way: “Progress on safeguards and model training has steadily driven down the first; the second—the theoretical blast radius—only grows as capabilities and access expand” (McGuinness and colleagues, Anthropic, May 2026). Their next sentence is the brief for this post: “The engineering question becomes how to cap the blast radius.”
The radius comes from what the agent holds, and the tool list is only part of that. In April 2026 an AI agent deleted a production database volume, and the infrastructure provider’s own account says how: the agent found one of the provider’s API tokens “stored locally on the user’s machine” and used it to call the delete operation “on a production volume” (Abdelwahab, Railway, 2026). The token “was provisioned with account scoped access, the maximum access possible.”
On my reading of that account, the deletion did not go through any deletion tool the agent had been given; the provider says it has since recovered the database. The process could read a file, and the file held a key.
So the rows of a blast-radius sheet have three origins: the credentials you hand the agent, whatever else its process can read or reach, and anything its actions can trigger that runs with wider access. The post on AI agent guardrails tells the rest of that incident and lists the checks that refuse a bad call. This post is about the size of what is left.
How do you estimate a blast radius before granting access?
Write one row per grant and compute a worst case for each. It is the smaller of two numbers: everything the credential reaches, or the calls the agent can make before it is stopped multiplied by the objects one call can touch. The formula is this post’s own, and it is deliberately simple enough to redo on paper.
For one row, the window is the minutes until someone notices plus the minutes it then takes to halt the agent. The calls that fit are one, plus the rate limit times the window, rounded down, and never more than a hard cap if one exists. The worst case is those calls times the objects each can touch, and never more than the credential’s reach.
Three rules keep the number honest. A rate with no code enforcing it is written as “none.” A time that was never measured is written as “unknown.” With no rate or an unknown window, the worst case is the whole reach, unless a cap exists that does not reset inside the window; a cap per run or per day counts once for every run or day the window can hold.
Each term has a lever, and the title’s three are among them. The mapping is mine; the dials are Chapter 17’s (“what the agent is permitted, where it runs, and which of its actions wait for a human”), and the stop is Chapter 13’s.
| Term in the estimate | What sets it | Lever |
|---|---|---|
| Which rows exist, and each row’s reach | The credential’s scope and lifetime | Least privilege |
| Rows nobody listed | What the agent’s process can read and connect to | Sandbox |
| Rate, cap, objects per call | Limits in the tool or a gateway | Limits in code |
| Minutes to notice | An alert that pages a person | Detection |
| Minutes to halt | A stop path somebody has drilled | Kill switch |
| Whether the worst case can be undone | A tested restore the agent cannot destroy | Undo |
| Whether a call waits for a person | An approval gate triggered by code | Approval |
Limits in code and approval belong to the guardrails post, and the consequence tier classifier sorts which actions wait for a person. Rate limits earn their place on this sheet for a reason OWASP’s entry on excessive agency gives: they “reduce the number of undesirable actions that can take place within a given time period, increasing the opportunity to discover undesirable actions through monitoring before significant damage can occur” (OWASP, LLM06:2025).
Why does the first call always count?
Because nobody reacts in zero minutes. The “one, plus” in the formula says that the first call lands whatever the window is, and for some rows the first call is the whole worst case: a bulk export, a send to an entire list, a deletion. The volume deletion above was a single request.
That gives the most useful distinction on the sheet. A stop bounds damage that accumulates: restarts, refunds, messages sent one at a time, money spent per hour. It does nothing for a row where one call does everything. Those rows have to be made safe by removing the grant, shrinking what one call can touch, making the action recoverable, or putting a signature in front of it.
Which lever shrinks which kind of access?
It depends on what inflates the row: reach, a single large call, or time. The table lists eleven common kinds of access with the worst single action for each and the lever I would pull first. The buttons filter by that lever; the judgments in the last two columns are mine.
| Kind of access | Worst single action | What inflates the row | What shrinks it | First lever |
|---|---|---|---|---|
| Read on a production database | Copy every table the role can read and send it out through any outbound path | Reach; one call is enough | A role limited to named tables or views; no outbound path from the process | least privilege, sandbox |
| Write on a production database | An update or delete that matches every row | Objects per call | A narrow tool for one operation in place of a connection; a row cap per statement; a point-in-time restore that has been tested | least privilege, undo |
| Shell or code execution | Anything the process’s user and network can reach, including credentials left on disk | Rows nobody listed | An isolated room with no secrets mounted and outbound traffic denied by default | sandbox |
| Cloud or infrastructure API | Delete a resource; scale to the account’s quota | Reach; spend that accrues by the hour | A role per verb and resource; a quota for that role; an expiry on what it creates; soft delete | least privilege, undo |
| Outbound email or messages | One send to a whole list | Objects per call; visible to third parties | Recipients set by code; a recipient cap per call; a signature for anything sent to a list | approval, least privilege |
| Payments and refunds | Many small payouts, each under the amount cap | Rate times window | An amount cap, a daily quota counted per credential, an alert on the count, a drilled stop | kill switch, approval |
| Secrets store, or credentials on disk | Read a broader credential, which adds every row of that credential to the sheet | Reach into other systems | Nothing mounted; short-lived tokens issued per task | sandbox, least privilege |
| Deploy pipeline or other automation | Trigger a job that runs with the pipeline’s wider credential | Reach by proxy | A separate permission to trigger; protected branches; a signature for production | approval, least privilege |
| Other agents or queued work | Fan out tasks that keep running after the parent is stopped | Window; concurrent runs | A stop that covers queues and children; limits counted per credential | kill switch |
| Public pages and posts | Publish a wrong statement | Visible at once | Draft only; a person publishes | approval |
| A loop on a schedule | The same small wrong action every night | A window measured in days | An alert on the output and not only on errors; a budget per loop; a stop that covers the scheduler | kill switch |
What does least privilege change: scope and lifetime?
Least privilege removes rows from the sheet and shortens how long a credential is worth anything. The principle is older than the field.
Saltzer and Schroeder stated it in 1975: “Every program and every user of the system should operate using the least set of privileges necessary to complete the job. Primarily, this principle limits the damage that can result from an accident or error” (Proceedings of the IEEE, 1975). Accident comes first in that sentence, before any attacker.
Chapter 17 adds the lifetime: “grant the minimum access the task requires, and nothing on standing. Cut the deputy’s keys per errand.” The chapter’s examples are all scope: “A credential scoped to one project beats an account-wide token; a tool that queries one table beats a connection string with admin rights; an agent that drafts email for your approval beats one that sends.”
Scope sets the reach of every row and decides whether the irreversible rows exist at all. Lifetime is a stop that works when nobody pulls anything: a token that expires in fifteen minutes is a fifteen-minute window for whoever copies it. One vendor describes the scope and revocation side of its own design: “the VM gets a per-session scoped-down token, and that token can be revoked independently of the user’s” (McGuinness and colleagues, 2026).
Read-only access is the largest single reduction, and it is not zero. The same write-up notes that “An agent with read-only DB access, for instance, can be deployed far more broadly than one that writes to prod.” A read is still a leak row if the process has any outbound path, which is the combination the lethal trifecta audit checks for and the post on the lethal trifecta explains.
What if the platform only issues broad tokens?
Then keep the broad token out of the agent’s process and put a narrow door in front of it. A practitioner put the complaint precisely in a February 2026 forum question: “If I want an agent to read logs or inspect env vars, I have to give it a token that also allows it to modify or delete things” (Hacker News, 2026).
The answer that fits the sheet is a small gateway that holds the broad token, exposes only the operations the task needs, and applies the rate limit. The agent’s row is then the gateway’s list, provided the token is somewhere the agent’s process cannot read. If the token sits in an environment variable of that process, the row is the token’s full scope, whatever the gateway offers.
What does a sandbox change on the sheet?
A sandbox deletes the rows that come from the process: files on the host, tokens left on disk, other sessions, open network access. Chapter 17 calls the logic subtraction: “If the credentials file is never mounted inside the boundary, no injection can read it: there is nothing to outwit, because the target is absent.”
It leaves the rows you granted on purpose exactly as wide as they were. The book’s warning applies to every destination you allow: “an allow-list entry is a capability grant, not a destination filter.” For the sheet, that means each allowed endpoint is a row of its own, with a worst single call.
For fleets, Chapter 20, Deploying and Scaling (in the full book) states the default in the section “Sandboxed Execution at Scale”: “credentials injected per task and scoped to it, nothing standing.” The mechanics, the tests that prove a wall holds and a policy to copy are in the guide to how to sandbox AI agent tool execution. I won’t repeat them here.
What makes an AI agent kill switch a control?
Three things: a named person who may pull it, a halt time measured in a drill, and a known answer to what state a stopped run leaves behind. Chapter 13, Writing the Outer Loop (in the full book) defines the kill switch as “one obvious, fast, tested way to stop every running loop and agent at once,” and adds: “A kill switch discovered broken during the fire is a design document for the next system.”
Public frameworks name the same parts, in their own words. The NIST AI Risk Management Framework (2023), under Manage 2.4, asks for mechanisms to “supersede, disengage, or deactivate” a system and adds that “responsibilities are assigned and understood” (NIST, 2023). The EU’s AI Act of 2024, for the systems it classes as high-risk, asks that the people assigned to oversee one be enabled to interrupt it through a stop button or “a similar procedure that allows the system to come to a halt in a safe state” (Article 14). The phrase “safe state” is the third item.
What does a missing stop procedure look like?
The clearest primary record I know predates agents. In 2012 an automated order router at a trading firm, plain software with no model in it, began sending erroneous orders. The regulator’s order records that it “routed millions of orders into the market over a 45-minute period, and obtained over 4 million executions in 154 stocks,” and that the firm “lost over $460 million from these unwanted positions” (SEC, Release No. 34-70694, 2013).
Read it for the window. Before the market opened, an internal system had sent 97 emails that referenced the router and an error; the firm “did not design these types of messages to be system alerts,” and staff “generally did not review them.” The firm “did not have supervisory procedures concerning incident response,” and one attempted fix “worsened the problem.” One of the order’s findings is a sentence for any runbook: the firm “needed clear guidance for its technology personnel as to when to disconnect a malfunctioning system from the market.”
I am not arguing that agents behave like that router. The record shows what the notice and halt terms look like when nobody has written them down: signals that page no one, and no rule for who disconnects what.
How do you test an AI agent kill switch?
Pull it on a running agent, on staging, and time it. The guardrails post explains why the switch should sit outside the agent’s own process. The drill below is what turns it into a number for the sheet.
- Someone other than the switch’s author pulls it, using only the runbook.
- The pull happens mid-run, with a write in flight; the time from the decision to the last side effect is recorded.
- After the pull, a direct call with the agent’s credential is refused, which shows the stop reaches the credential or the gateway and not only the loop.
- Queued runs, scheduled runs, retries and child agents do not start.
- The agent’s process is made to hang, and the stop still works.
- The outcome of the call that was in flight is known: completed, rolled back, or needing repair by a person.
- Everything the run left behind is listed and removed: scaled resources, locks, temporary credentials, half-applied changes.
- A seeded fault triggers the alert that should lead to the pull, and the time from fault to page is recorded.
- The run resumes from its record without repeating a write.
- Two named people can pull the stop outside working hours, and both have the access to do it.
- The date, the notice time and the halt time go on the sheet; the drill is repeated after any change to the credential, the orchestrator or the model.
What does stopping leave behind?
Whatever the run had already created, and a decision about resuming. Halting an agent does not remove the instances it started or recall the message it sent, so each row of the sheet names the state left behind and who removes it.
Chapter 12, Oversight and Autonomy (in the full book) gives the design rule for the other half: “interruption is only cheap if resuming is.” A stop that forces a restart from zero is a stop people hesitate to pull. A run that keeps a checkpoint and uses idempotent tools and safe retries can be stopped early and often.
An automatic stop shortens the notice term to seconds for the faults it knows. The circuit breaker simulator shows one kind. It complements the manual switch and does not replace the drill.
Can this AI agent have production access?
Answer row by row, with three questions; the blast radius of AI agents is approved one grant at a time. The rule is this post’s own. Access is granted per row and never per agent, and any “no” means the row is removed or fixed before the credential is issued.
- Bounded? The worst case is a number, and a named mechanism outside the model enforces it: a scope, a cap, a quota.
- Recoverable, signed or accepted? Either a restore has been tested and the agent’s credential cannot destroy the copy, or every action waits for a signature, or a named person has accepted the bounded loss in writing.
- Stoppable? Something ends the damage in a time that was measured: a drilled stop with two named people who can pull it, or an expiry on what the action creates.
A row where one call does the whole worst case cannot be helped by the third question, so it has to pass on the first two.
BLAST-RADIUS SHEET: [agent], [owner], [date], [requested by]
A. Where the rows come from (list all three, then write one row each)
- Credentials the agent is given: [name, scope, lifetime]
- Anything else its process can read or reach: [files, tokens on
disk, network destinations, other sessions]
- Things its actions can trigger with wider access: [pipelines,
queues, other agents]
- Concurrent runs sharing these credentials: [n]
B. One row per grant
Grant: [an action, a dataset, a budget, or another system]
Worst single call: [...]
Reach: [how many objects the credential can touch] [unit]
Objects per call: [1 | n | everything in reach]
Rate limit: [n per minute | none], enforced in: [...],
counted per: [run | credential]
Hard cap: [n per run or per day | none], enforced in: [...]
Notice: [n minutes], by: [alert that pages a person | "a customer
tells us"], last fired in a test on: [date]
Halt: [n minutes | unknown if never drilled], by: [stop path],
last drilled on: [date]
Worst case = the smaller of reach and (calls x objects per call)
calls = 1 + rate x (notice + halt), rounded down, at most the cap,
times the runs that can start inside the window (concurrent or
scheduled) if the limit or the cap is counted per run
no rate limit or an unknown time: the whole reach, unless a cap
exists that does not reset per run or per day
=> [number] [unit]
If it costs money over time: [objects] x [cost per hour] x [hours
until removed] = [...]
If each object carries an amount: [objects] x [largest amount
each] = [...]
State left behind after a stop: [...], removed by: [...]
C. Three questions per row (any "no": remove the grant or fix the row)
1. Bounded? The worst case is a number, enforced by:
[mechanism outside the model] [yes | no]
2. Recoverable, signed or accepted?
- restore tested on [date], and the agent's credential cannot
destroy the copy; or
- every action waits for a signature; or
- [name] accepted the bounded loss in writing on [date]
[yes | no]
3. Stoppable? Notice and halt measured in a drill on [date], and
[name] and [name] can pull the stop; or what the action creates
expires after [n hours] [yes | no]
(If one call does the whole worst case, the stop cannot help:
the row must pass on 1 and 2.)
D. Verdict
Granted rows: [...] Removed rows: [...] Open rows: [...]
One sentence for the board: "On its worst day, before we stop it,
this agent can [...]. It cannot [...]."
Review again on: [date], and on any change of credential, tool,
orchestrator or model.
What does the sheet say for one agent?
Here is a worked example with illustrative numbers, computed by a short script. A team that owns 6 of a company’s 40 services wants an on-call triage agent that reads alerts, restarts a stuck service, adds capacity under load and drafts status updates.
The request as written is “give it what the on-call engineer has”: a standing cluster-admin credential, a cloud key and the status page’s key. A restart storm would page someone in 5 minutes. The stop has never been drilled.
| Row | As requested | Worst case as requested | After the levers | Worst case after |
|---|---|---|---|---|
| Restart a service | Admin credential, 40 services, no rate limit | 40 of 40 services, repeatedly | Role limited to restart on 6 services; one restart per 2 minutes, counted per credential; cap of 6 per run; notice 5 min on an alert seeded at this limited rate, halt 2 min, drilled | 4 restarts, so at most 4 of 6 services |
| Delete a volume | Same credential, 12 volumes | 12 volumes, one call each, back to the last backup | The role has no delete verb | 0 |
| Read secrets | Same credential, 85 secrets | 85 secrets to rotate | The role cannot read secrets; the process holds one 15-minute token | 0 |
| Add capacity | Cloud key, 200 instances of quota headroom | 200 × $0.50 × 24 h = $2,400, found on the next day’s bill | A cap of 4 instance launches per day for the role, counted per credential; each instance expires after 2 hours | 4 × $0.50 × 2 h = $4 a day |
| Post to the status page | Publishing key | Unbounded public posts | Draft only; a person publishes | 1 per signature |
The restart arithmetic: the window is 5 + 2 = 7 minutes, the rate is 0.5 per minute, so 1 + 3 = 4 calls, under the cap of 6. The $0.50 per instance-hour is a round illustrative figure and not anyone’s price. Two rows that could not be undone are gone, the restart row fell from 40 services to 4, and the spend row fell from $2,400 to $4, both per day.
Now the three questions.
The restart row is bounded by the role and the gateway’s limit, recoverable because a restarted service comes back, and stoppable in a drilled 7 minutes. The capacity row is bounded by the daily launch cap; money spent is not recoverable, so a named person accepts $4 a day in writing; the expiry ends it, since halting the agent would leave the instances running. The status row is signed. All three granted rows pass.
From the table, the sentence for the board writes itself: “On its worst day, before we stop it, this agent can restart four of the six services it looks after, start four extra instances a day that run two hours each, and publish nothing without a signature. It cannot delete a volume or read a secret.”
What moves the numbers?
The stop, where the limit is counted, and the token’s lifetime. The same script gives these variations on the restart row.
| Change to the “after” restart row | Worst case |
|---|---|
| As drilled: halt in 2 minutes | 4 of 6 services |
| Halt never drilled, so the window is unknown | 6 of 6: the scope holds; the cap of 6 per run holds only if no second run starts |
| Halt measured at 30 minutes | 6 of 6: the window is 35 minutes, so 1 + 17 = 18 calls, cut by the cap and the scope |
| Five concurrent runs, limit counted per run | 6 of 6: 4 × 5 = 20 calls against 6 services |
| Five concurrent runs, limit counted per credential | 4 of 6 |
| The 15-minute token is copied and nobody revokes it | 6 of 6: 1 + 7 = 8 restarts before it expires |
| The same token with no expiry | 6 of 6, repeatedly |
Two lines deserve a second look. An undrilled stop costs the row a third of the team’s services here, and it would cost far more on a row with a wider scope. A limit counted per run does not bound a fleet.
Chapter 20’s instruction, in “Rate Limits, Queuing, and Backpressure,” is to “Cap each lane, each tenant, and each agent” so that “one runaway workload, one looping agent, one enthusiastic customer cannot starve the rest of the fleet.” Count limits where the credential is checked.
You don’t have to grant every row on the first day. Start with the read rows, add a write row in staging, then production with the narrow scope. The rollout plan generator drafts that sequence of rungs with a way back at each one.
What does the estimate leave out?
An estimate of the blast radius of AI agents leaves out probability, correctness inside the limits, and any proof that the limits are built right. It is an upper bound on the cost half of the bet. It says nothing about how often the agent will be wrong, and I know of no public data that would let you put a defensible likelihood on a row.
It also counts nothing for a wrong action that stays inside every limit. A read-only analytics agent that writes a wrong query hands an executive a wrong number, and no row on the sheet shows it. One commenter, in a reply to a comment describing exactly that setup, asked, “How do you validate that the reports are correct?” (Hacker News, April 2026). That is a question for evaluation, and access limits do not answer it.
Four more gaps remain. The sheet assumes each mechanism works as written, which only a negative test shows. Chains are hard to list: a row that triggers a pipeline inherits the pipeline’s rows, and you will miss some.
Units do not add up across rows, so there is no single score, and I think a single score would mislead. And the acceptance in question 2 is a judgment about what loss your company can carry, which the sheet records and cannot make.
The cost is real as well. Scoped roles, a gateway and a drill are days of work, and narrow access means more escalations to a person. For an agent that only drafts text for someone to read, a short sheet with two or three rows is enough.
The sentence you sign
The blast radius of AI agents is a number per grant that you can write before the first credential is issued, and the approval you give is a signature under that number. The policy that says who fills in the sheet and how often it is reviewed belongs in an AI agent governance framework. The same sheet works from the buyer’s side: when you weigh a vendor’s claims about AI agent reliability for enterprise use, ask for its rows, its worst cases and the date of its last stop drill.
If the request on your desk says “give it what the engineer has,” send back the sheet and ask for one row per grant. Chapter 17 puts the reason in one line: “You cannot buy a model that is never wrong; you can always shrink what wrong costs.”
Chapter 17, “Security, Safety, and Guardrails,” develops the three dials (in the full book), Chapter 13, Writing the Outer Loop (in the full book) covers the kill switch and the brakes for unattended loops, and Chapter 20, Deploying and Scaling (in the full book) covers per-task credentials and caps at fleet scale. The AI agent security guide places this post beside its neighbors, or you can see the formats.
Questions readers ask
- What is the blast radius of an AI agent?
- The book defines blast radius as everything a step, a run, or an agent could break or leak if it went as wrong as possible. It is set by the credentials the agent holds and by what its process can reach, so it can be listed and estimated before the agent runs. How likely a failure is does not enter it.
- Can AI agents be trusted with production access?
- The useful form of the question is per grant, not per agent. Grant a row of access when its worst case is a number enforced outside the model, when that worst case can be undone or each action waits for a signature or a named person has accepted the loss, and when a drilled stop that two named people can pull, or an expiry, ends it in a measured time (a row where one call does everything has to pass on the first two). Remove or fix any row that fails.
- What is an AI agent kill switch, and how do you know it works?
- The book defines a kill switch as one obvious, fast, tested way to stop every running loop and agent at once. It works if a drill shows it: someone other than its author pulls it mid-run from the runbook, a direct call with the agent's credential is then refused, queued work does not start, and the time from decision to last side effect is written down with a date.
- Is read-only access safe for an AI agent?
- It is much narrower, and it is not zero. A read is a leak row if the agent has any way to send what it read somewhere else, and the worst case is every table or file the credential can read. A wrong answer built from a correct read is a separate risk that no access limit covers.
- Does a sandbox reduce the blast radius?
- Yes, for one part of it. A sandbox removes the rows that come from the agent's process, such as files on the host, credentials left on disk and open network access. It does not narrow what an allowed tool or an allowed destination can do, which is the job of the credential's scope and of limits in code.
Sources
- McGuinness, Grace, De Jonghe, Eaton, Ribbink (Anthropic) (2026). How we contain Claude across products
- Jerome H. Saltzer, Michael D. Schroeder (1975). The Protection of Information in Computer Systems (Proceedings of the IEEE 63(9))
- OWASP Gen AI Security Project (2025). LLM06:2025 Excessive Agency
- Mahmoud Abdelwahab, Railway (2026). Your AI wants to nuke your database. Guardrails fix that
- U.S. Securities and Exchange Commission (2013). In the Matter of Knight Capital Americas LLC, Release No. 34-70694 (order of 16 October 2013)
- European Union (2024). Regulation (EU) 2024/1689 (AI Act), Article 14: Human Oversight
- National Institute of Standards and Technology (2023). AI Risk Management Framework 1.0, Core (Manage 2.4)
- NBenkovich (Hacker News) (2026). Ask HN: How do you give AI agents access without over-permissioning?