Home / Blog / Security, reliability and cost / AI Agent Guardrails: A Deterministic Layer the …

Security, reliability and cost

AI Agent Guardrails: A Deterministic Layer the Prompt Can't Be

AI agent guardrails are checks in code that decide if an action happens. See 19 candidates sorted by where each runs, its cost and its test. Start here.

By Enrique Gutiérrez · Published · 27 min read

AI agent guardrails are checks that run in code around the model and decide whether an action happens: a cap inside a tool, a scoped credential, a budget on the loop, an approval gate. A rule in the prompt is a request the model usually honors. One question sorts every candidate: can text argue with it?

I wrote this for a backend engineer holding a ticket that says “add guardrails to the agent.” The AI agent guardrails in this post have no attacker to catch. The model is wrong once in some number of runs, a task drifts, or a loop won’t end, and the question is which code refuses, caps or stops it and how you know that code works. The attacker’s side of the story is in the companion post on how to prevent prompt injection in AI agents.

Does a rule in the system prompt count as a guardrail?

No. A rule in the system prompt is text the model reads and weighs on every pass, and no code halts the run when the weighing goes the wrong way. Chapter 17 of the book (in the full book), in the section “Guardrails as a Separate, Deterministic Layer,” asks what enforces three such sentences and answers: “There is no interpreter that halts on violation, no permission system consulting your list.”

The same paragraph lists how the request fails, and two of the three ways need no adversary. The model can “drift from it under a hundred turns of accumulated context, or simply draw an unlucky sample on an ordinary afternoon.”

One public case shows the shape. According to a press report of a user’s own posts (The Register, July 2025), a coding service deleted his production database “despite his instructions not to change any code without permission.” Later, after trying to have the service freeze code changes, he wrote that there was “no way to enforce a code freeze” in apps of that kind. He also wrote, “you can’t not separate preview and staging and production cleanly.”

Keep the rule in the prompt, because it is free to write and can lower how often the model tries. Then apply the sorting question, which is this post’s own formulation: when the rule is violated, which code refuses or halts, and can any text the model reads change that? A freeze that is a sentence fails the test. A freeze that removes the write tools from the run passes it.

What does “AI agent guardrails” mean, and to whom?

The phrase “AI agent guardrails” has no standard meaning, so ask what the speaker includes. The book’s definition is the one this post uses, and it carries a tension worth reading closely: “A guardrail is the field’s answer: a check that lives outside the model’s reasoning and runs deterministically, whatever the model has decided.” The next sentence widens it: “Concretely it is ordinary code—or, as we will see, occasionally a second, separate model—stationed at the agent’s boundaries, examining what passes and empowered to allow it, block it, modify it, or escalate it to a person.”

I read the two sentences as naming two properties, and the chapter supports the reading a few paragraphs on, where it says of a model-based guard that “its verdicts are probabilistic.” What runs deterministically is the checkpoint: it sits outside the model, it always executes, and its verdict binds. Whether that verdict is certain depends on who renders it: code, a model or a person. A guardrail in this post means the first property, and the inventory below marks each row with the second.

Other sources draw the line elsewhere. One lab’s guide to building agents (2025) describes an agent’s instructions as “Explicit guidelines and guardrails defining how the agent behaves,” and it places authorization outside the word: guardrails should be coupled with “authentication and authorization protocols, strict access controls, and standard software security measures.” The OWASP entry on excessive agency (LLM06:2025) doesn’t use the word in its guidance and requires the control: “Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not.”

Is a guardrail library or product enough?

For an agent, no. A product that inspects text going into and out of the model covers the layer the book ranks below the tools: “Prose is recoverable; an apology retracts a bad sentence.” The destructive call in the incident below was authenticated and looked ordinary, and I read nothing in the write-up that a content check would have flagged. When a vendor says “guardrails,” ask whether authorization, tool limits and budgets are inside the word.

Which AI agent guardrails exist, and can text argue with each?

The table lists nineteen candidate AI agent guardrails with where each is enforced, whether text can argue with it, what it costs and how to test it. Two rows are on it to be classified: the system-prompt rule, which is a request, and the log, which detects and stops nothing.

In the fourth column, which applies the second property, “No” means code renders the verdict and no sequence of tokens can change it; it still depends on the check being built, configured and wired in correctly, like any access control. The cost and test columns are my judgment except where a source and year are named.

The last column says which of five harms the row is aimed at, and the buttons filter on it; the two rows that are there to be classified carry all five tags, so they stay visible under every filter. A wrong action is a call outside what the task or user authorized. An unsafe write is a destructive, irreversible or externally visible change. A data leak is data leaving its boundary.

Runaway cost is a loop, retry storm or fan-out that spends without converging. Bad output is text that is malformed, off-policy or unsupported. Injection is deliberately absent from the list: it is a cause that can produce any of the five, and the checks here don’t need to know the cause.

Guardrail What it stops Where it is enforced Can text argue with it? What it costs How to test it Aimed at
Typed tool with limits in code (amount cap, allowed recipients, allowed tables) Out-of-policy arguments, whatever the model decided Inside the tool function, next to the side effect No Engineering per tool; some legitimate edge cases refused Call the tool directly with an out-of-range argument; expect a refusal wrong action, unsafe write
Tool allow-list per task, or the tool removed Calls to tools the task doesn’t need Harness: the tool list sent to the model, and the dispatcher No; structural, if code picks the list A less capable agent; more escalations Assert the dispatcher rejects an unlisted tool name wrong action, unsafe write, data leak
Scoped, short-lived credential per task Everything outside the credential’s scope The downstream system’s own authorization No Design time; provisioning friction With the agent’s credential, attempt the forbidden operation by hand; expect a denial unsafe write, data leak
Identity arguments set by code from the session (tenant, user, account, environment) Cross-tenant, cross-user and cross-environment reads and writes Application code and the database’s row-level policy No Engineering per tool Supply another tenant’s id in a model-filled field; expect it ignored or refused wrong action, data leak
Approval gate keyed to consequence tier The irreversible or high-stakes actions a person reads Harness, between tool selection and tool execution Not the trigger. The verdict is a person’s: they tire, and they read text the model may have written Latency and attention. About 93% of prompts approved in one vendor’s telemetry (2026); one threat in three missed in a public game (2026) Assert a tier 3 or 4 call is held until a person decides, and that a changed argument voids the approval unsafe write, wrong action
Undo window, soft delete, staged change Permanence of a destructive action, for as long as the window lasts and only if the same credential cannot delete the undo copy The downstream system No Storage; delayed reclamation Delete through the agent’s path; restore within the window unsafe write
Budget on steps, tokens, wall-clock time and money, with a timeout per call A loop that never converges; one pathological tool result; a call that never returns Harness loop, checked before each model call; the timeout in the tool wrapper No Legitimate long tasks cut off; needs tuning from traces Stub a model that always requests the same tool, and a tool that hangs; assert the run stops at the cap and says so runaway cost
Rate limit and quota per tool, per run and per user Bursts of repeated actions; fan-out Tool wrapper or gateway, and the downstream API No Legitimate bursts throttled Issue one call more than the quota; expect it refused runaway cost, unsafe write
Circuit breaker on a failing tool, with one retry budget for the whole run Retry storms; timeouts paid on a dead dependency Tool wrapper and harness No Code to maintain; a breaker that counts “not found” as failure cuts off a healthy tool Make the dependency fail; assert the breaker opens, then half-opens runaway cost
Kill switch, plus an automatic stop after repeated guard trips The next step of a run, or of every run; calls already dispatched continue unless the downstream is cut too Outside the agent’s process: orchestrator, gateway or credential revocation No, once pulled; the manual switch waits for a person or an alert Stopping is expensive unless resuming was built Trip it on a live staging run; measure time to halt; check nothing is half-applied runaway cost, wrong action, unsafe write
Sandbox with default-deny egress and no secrets mounted Reach beyond the boundary; an allow-listed endpoint is inside it Operating system, container runtime or hypervisor No; structural Infrastructure; each new capability needs configuring From inside, read a host secret and reach an unlisted host; expect both to fail data leak, unsafe write
Output schema validation, with a canary field that must stay empty Malformed output; content the model flags itself After generation, before the next system No for shape; yes for the canary field, which the model has to fill One more inference per repair; cap the retries Feed malformed and canary-filled outputs to the validator; expect rejection bad output
Output redaction by pattern (card numbers, keys, internal hostnames) Secrets in a known format After generation, before the user or the log No, within the formats it knows Little latency; some false redactions Unit-test each pattern with strings that must and must not match data leak, bad output
Grounding check (cited claims appear in the supplied context) Quotations and citations that are not in the supplied context After generation No for exact quotes; yes if a model judges A string match is cheap; a judge is one more inference Labeled grounded and ungrounded answers; measure both error rates bad output
Input limits: length, format, scope, size of what a tool feeds back Oversized or out-of-scope requests before they cost a model call Before the model call No for length, format and size; yes if a classifier decides scope Some legitimate edge cases rejected Boundary-value tests runaway cost, bad output
Trained input or output classifier Known categories of off-policy content, at volume Before or after the model Yes; probabilistic For one vendor’s classifiers (2025): refusals up 0.38 points on 5,000 sampled production conversations, 23.7% more inference compute Labeled set; report false-block and miss rates with dates; rerun on every model change bad output, data leak
A second model judging each tool call Some out-of-scope or overeager actions Harness, before tool execution Yes; probabilistic One more inference per judged call. For one vendor’s action classifier (2026): 0.4% false blocks on 10,000 real tool calls, 17% of 52 real overeager actions missed The same, on recorded tool calls wrong action, unsafe write
Rule in the system prompt Nothing by itself; it can lower how often the model tries Inside the model’s context Yes; it is a request Free to write; competes with every other instruction Can be sampled, never proven wrong action, unsafe write, data leak, runaway cost, bad output
Log of every tool call and guard trip, with alerts Nothing; it shortens the time a failure stays unseen Harness and downstream systems Not applicable; detective Storage; alert tuning Trip a guard on purpose; check the alert arrives and names the call wrong action, unsafe write, data leak, runaway cost, bad output

Which of your backend instincts carry over?

Most of them carry over, with one change that repeats down the list: the caller you are constraining is a model whose requests are plausible and occasionally wrong, so every check has to work without the caller’s cooperation. The mapping is mine.

You already do this Agent equivalent What is different
Authorization at the resource, never in the client The downstream system and the tool check the caller’s scope The client is the model; its opinion about what it may do is never consulted
Validate every request body Validate every tool call’s arguments in the tool, and model output before the next system uses it Arguments are always well-formed and sometimes wrong, so validate meaning (owner, range, tenant) as well as shape
Rate limits and quotas Per-tool quotas, plus a budget per run on steps, tokens, time and money One user request can fan out into hundreds of calls, so the limit attaches to the run as well as the user
Circuit breaker around a flaky dependency A breaker around each external tool, and a stop after repeated guard trips The second kind trips on refusals, which are evidence about the run and say nothing about the dependency’s health
Feature flag or kill switch An operator stop outside the agent’s process A stopped run has state; stopping is cheap only if resuming was built
Idempotency keys An approval bound to the exact action and arguments; safe retry of a write The retry may be the model asking again in different words
Audit log A log of every tool call and every guard trip A refusal recorded as an ordinary tool result is invisible in the trace

Where does each check go?

Start with what the agent can reach at all, then put each check as close to the side effect as it will go: five places, in order. The ordering is this post’s synthesis; the book supplies the principle, “put the check next to the side effect,” and OWASP supplies the second step.

  1. The process boundary. What the agent’s process can reach at all: a sandbox with no production credentials inside, a gateway in front of its calls, and the operator’s stop.

  2. The downstream system’s own authorization. A scoped credential, a row-level policy and an undo window hold for every caller, including an agent that goes around your code.

  3. The tool. Argument limits, identity arguments filled from the session (tenant, user, account, environment), quotas and a timeout live in the function that causes the side effect.

  4. The harness. The harness (the code around the model that runs the loop) owns the tool allow-list, the approval gate, the budgets and the trip counter. Checks on what goes into the model and what comes out run here too.

  5. A model or classifier, last. Use one where code can’t decide the question, never as the only check in front of a side effect, and make it finish before the action starts. One agent framework’s documentation (undated; read October 2026) says of an input guard that runs in parallel with the agent: “the agent may have already consumed tokens and executed tools before being cancelled.”

What happens when the downstream check is missing?

The agent’s mistake executes at once and nothing in the path can take it back, and one documented incident shows how. In April 2026 an AI agent deleted a production database volume, and the infrastructure provider’s own write-up says what the agent did: it found one of the provider’s API tokens “stored locally on the user’s machine” and called the delete operation directly. The write-up says “The request was authenticated, and our API honored it the same way it would for a CLI command or a CI pipeline,” and that the agent “wasn’t told to delete a database. It decided that deletion was a reasonable step toward fixing something unrelated.”

The provider names two missing controls. The token “was provisioned with account scoped access, the maximum access possible,” and the delete call ran “immediately, with no way to undo it” while the dashboard gave the same action a 48-hour window. The call also “performed a cascading delete on the model, making the backups look unavailable in the UI.” Its first fix was a soft delete for 48 hours on every delete; it also delayed the deletion of backups and said it would look again at how token scopes are presented.

By its own account the provider has since recovered the database; the customer’s public account differs on the details of the recovery, so I have used only what the write-up states. On my reading of that account, the process boundary was the only place on the customer’s side that would have been in the path, because the agent used a token it found on disk and never went through a tool. The provider’s summary generalizes: “make the destructive thing slow, make the recoverable thing fast, and put the actual point of no return as far away from a single click as possible.”

How do you stop an agent deleting, sending or spending?

Combine three things: limits inside each write tool, an approval gate on the few actions that are irreversible or high-stakes, and budgets with a stop path for runs that won’t end. The first is the top four rows of the inventory: limits in the tool, the tool list, the credential and the identity arguments. The other two follow.

How do you add approval gates to AI agents without wearing them out?

Gate by what the action costs when it is wrong. Chapter 12, Oversight and Autonomy (in the full book), in “Approval Gates and Escalation,” sorts everything an agent can do into “consequence tiers—read-only, reversible, externally visible, irreversible” and assigns oversight in four sentences: “Read-only actions run without ceremony; gating them buys no safety and teaches the person approving them to stop reading. Reversible actions run freely too, provided the system logs enough to undo and audit them. Externally visible actions, anything a third party will see, deserve a review queue. Irreversible or high-stakes actions wait for a signature, every time.”

An approval gate is a guardrail only if its trigger is code: “the gate must live in your code, never in the agent’s judgment.” Keying approval gates to consequence has a corollary teams skip: “the agent’s confidence plays no part in it.”

What does the classifier show for six actions?

It shows three signatures, one queued review and two actions that run. The classifier below opens on six actions from three agents: a support agent that looks up an order, drafts a reply, sends it and issues a refund; a coding agent that deletes files; and an agent that emails a customer list. The team’s current rule is “ask a human when the agent is less than 80% sure,” so no action is gated for what it is.

With JavaScript on, the Consequence tier classifier runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Consequence tier classifier on its own page to share a result by link.

issue_refund, delete_files and email_customer_list land in tier 4 (“signature, every time; sync block”) because I answered yes to the tool’s question “Does it move money, delete data, or touch the public?” I counted a send to a whole customer list as touching the public. One reply to the customer who wrote in, send_reply, is tier 3 and goes to a review queue; draft_reply is tier 2 and runs with a log; look_up_order is tier 1 and runs with no gate.

The classifier also raises three warnings: tier 4 actions run with no signature today, an externally visible action reaches a third party with no review, and the current rule mentions the model’s confidence.

Now rebuild delete_files so that it only moves version-controlled files into a restorable trash. The honest answers change to “no” on the second question and “yes, and it’s logged to undo” on the question about undoing the change, and the classifier moves it to tier 2.

The tool’s questions do not ask about availability: a reversible change that can take production down is high-stakes, and I would answer yes to the second question for it.

Why not ask for approval on everything?

Because approval is a guardrail whose weakness has been measured, and the vendor that measured it says attention falls as prompts multiply. One vendor reported that users of its coding agent “approved roughly 93% of permission prompts” (company telemetry, no sample size given; Anthropic, 2026). A public browser game in which players approve or deny an agent’s commands logged “over 40,000 runs and 409,000 individual approve/deny decisions,” and “The average player missed 1 in 3 threats,” a miss being an approved command the game counts as a threat (Wauters, August 2026). The author adds his own caveats: about 34% of the commands shown were threats, far above daily work, and players were on a clock.

Replacing the person with a model has its own numbers. One vendor’s classifier that approves tool calls in place of a human wrongly blocked 0.4% of 10,000 real tool calls from its own employees and let through 17% of 52 real overeager actions, and its authors write that it “is not a drop-in replacement for careful human review on high-stakes infrastructure” (Anthropic, March 2026). The vendor ran the study on its own product.

What budgets should a run have?

A run should have several at once: steps, tokens, wall-clock time and money. Chapter 3, The Agent Loop (in the full book) defines the term in “Stop Conditions, Budgets, and Frameworks”: “A budget is a hard ceiling enforced by the harness, indifferent to the model’s opinion.”

The loop has three ways out, and they are not equally trustworthy.
Figure 3.5 The loop has three ways out, and they are not equally trustworthy. A verified check—a suite that goes green, a file that parses—is evidence, and where one exists it should pronounce the run finished (in accent). The model’s own “done” is only testimony, believed with caution. And the budget cap is the safety net: it settles nothing about the work, but it guarantees the loop stops, so it must always be there. Reuse this diagram

Read the cap from your own traces; the agent cost per task estimator shows how re-read history and retries multiply the tokens in one run. When a budget fires, the instruction is “Fail loudly”: save the transcript and say the run was stopped.

A budget checked between model calls can’t end one tool call that hangs, so each call needs its own timeout. Chapter 18 says it in five words: “Give every call a deadline.”

What is an AI agent kill switch, and where does it live?

A kill switch is, in the words of Chapter 13, Writing the Outer Loop (in the full book), “one obvious, fast, tested way to stop every running loop and agent at once.” The book doesn’t say where the switch lives.

The placement is this post’s point: put it outside the agent’s own process. A stop that depends on the agent reading a message, or on a handler inside a wedged process, fails in the cases it exists for. An orchestrator that stops scheduling steps, a gateway that refuses the agent’s calls and revocation of its credential all work without the agent’s cooperation. Each stops the next step; a call already dispatched keeps running unless the downstream system is cut off too.

Pair the manual switch with an automatic one. The vendor whose action classifier is measured above stops a session after “3 consecutive denials or 20 total” (Anthropic, March 2026), which is a breaker on guard trips.

What does a guardrail spec for one tool look like?

It is one page per write tool that says what refuses a bad call, where that code lives and which test proves it. Paste it into the pull request and treat every empty line as work that remains.

GUARDRAIL SPEC: [tool name], [owner], [date]

1. What the tool may do
   - Changes: [what state, in which system]
   - Worst single call if every argument is wrong: [...]
   - Not offered at all: [operations left out of the tool]

2. Limits enforced in code (file and function for each)
   - Process boundary: [what the agent's process can reach; which
     credentials are absent from it]
   - Downstream: credential [name], scope [...], lifetime [...]
   - [argument]: [type, allowed range or allowed set]: enforced in [...]
   - Quota: [n] calls per run, [n] per user per day: enforced in [...]
   - Timeout per call: [n] seconds: enforced in [...]

3. Identity arguments set by code from the session, never by the model
   - [tenant / user / account / sender / environment]: read from [...]

4. Consequence tier and approval
   - Tier: [1 read-only | 2 reversible | 3 externally visible |
     4 irreversible or high-stakes]
   - If the tier depends on the arguments: the rule in code that
     decides, or "treated as its worst case"
   - Gate: [none | log for undo | review queue | signature every time]
   - The approval is bound to: [exact arguments shown to the approver];
     any change to them forces a fresh decision
   - What the approver sees is rendered by code from the arguments,
     never the model's summary
   - The approval is verified by: [the tool, checking a token bound to
     the arguments | the harness only]
   - Undo path and window: [...] or "none"; the agent's credential
     cannot delete the undo copy

5. Budgets this tool draws on
   - Run: [steps], [tokens], [wall-clock], [money]
   - Retries: [n] per call, only with idempotency key [how derived];
     counted against the run's single retry budget
   - Outcome unknown (timeout on a write): [read back | escalate];
     never resend without the key

6. What happens on a trip
   - Check refuses: [halt the run (default) | specific error returned
     to the model, for a limit the model can satisfy]
   - Check throws or times out: block. Never continue.
   - Model retries allowed after a fed-back refusal: [n], then halt
   - Run halts after [n] consecutive or [n] total guard trips

7. Stop path
   - Who can stop this tool, and how: [flag, gateway rule, credential
     revocation]; it lives outside the agent's process
   - A call in flight when the stop fires: [completes | rolls back |
     needs manual repair]
   - Time to halt, last measured: [...] on [date]

8. Tests that prove each guard (no model in the loop)
   - One negative case per limit in section 2: expect a refusal
   - Another identity's id in a model-filled field: expect it ignored
   - The check made to throw: expect the call blocked
   - A tool that hangs: expect the timeout to end the call
   - A stub model that repeats the call: expect the budget to stop the run
   - A tier 3 or 4 call: expect it held until a person decides, and a
     changed argument to void the approval
   - Undo: perform the action, then restore within the window
   - The production configuration loads every guard above

9. Model-based checks in front of this tool (not counted above)
   - [classifier or judge]: false-block rate [...], miss rate [...],
     measured on [n] labeled cases on [date], or "not measured"

10. Log fields: [tool, arguments with secrets and personal data
    redacted, prior value or undo handle, verdict, which guard
    tripped, approver]

How do three real actions come out?

Here are three write actions from the classifier, one of them below the signature tier, each walked through the inventory, the five places and the spec. The answers are illustrative.

Refund (support agent) Delete files (coding agent) Send a reply (support agent)
Process boundary and downstream system The process holds one payment credential, scoped to refunds on the merchant’s own orders Sandbox: only the workspace is mounted, no production tokens on disk; the trash sits outside what the tool can erase Mail credential scoped to one sender address; the provider’s own sending limits
In the tool Amount cap; the order must belong to the session’s customer; refunds per customer per day; timeout The path must resolve inside the workspace; files move to a restorable trash; files per call; timeout The recipient is set by code from the ticket; one reply per ticket; timeout
In the harness Signature gate; run budgets; trip counter Tool allow-list; signature gate while it is a raw delete; run budgets; trip counter Review queue; schema and redaction checks on the body; run budgets; trip counter
Tier and approval Tier 4: signature every time, bound to order and amount Tier 4 as a raw delete; tier 2 once it is a logged, restorable move Tier 3: review queue, approval bound to the exact body
Retries Only with the same idempotency key; unknown outcome: read back or escalate Safe to repeat: the file is already in the trash Only with the same idempotency key; unknown outcome: escalate
On a trip Over the cap: the refusal goes back to the model once. Another customer’s order: halt A path outside the workspace: halt The body fails schema or redaction: one repair, then halt
Test Direct call over the cap; another customer’s order id; the policy lookup made to throw; a changed amount voids the approval A traversal path refused; delete, then restore; read a host secret from inside A model-supplied recipient ignored; a canary field filled; an edited body voids the approval

Halting is the default. Chapter 3 lists “a guardrail trip” among the events where “the safe response is to halt it, loudly,” and the error table in this site’s post on idempotent tools and safe retries says the same. Chapter 17 allows feedback for “the milder trips (an output that failed groundedness, a malformed schema),” with capped retries, which covers the reply’s body.

Feeding one over-cap refusal back is my exception, for a limit the model can satisfy by asking for less. A trip here means a refusal by a limit, a policy check or an output check. Every trip counts toward the trip counter whether or not it was fed back, and a malformed argument is an ordinary tool error that counts for nothing.

What do AI agent guardrails cost?

AI agent guardrails that are ordinary code cost engineering time and some refused edge cases; model-based checks cost an inference each and bring two error rates; approval costs attention. The book doesn’t soften it: “Every check adds latency; every model-based check adds an entire additional inference, in both time and money, to every request it examines.”

For classifiers, a 2025 preprint from one vendor reports that its jailbreak classifiers raised refusals by 0.38 percentage points on a random sample of 5,000 of its own production conversations and added 23.7% inference overhead (Sharma and colleagues, 2025).

False blocks compound over a run, and this arithmetic is mine and illustrative. If each of k guarded steps is falsely blocked with probability f, and the steps are independent, a clean run finishes untouched with probability (1 − f)k. At f = 0.4% and 30 steps that is about 89%; at 8.5%, the false-block rate that vendor’s 2026 evaluation reports for its first stage alone, about 7%. A guard on every step needs a far lower false-block rate than a guard on one chat turn.

Fail-closed has a price too. The book’s rule is “The failure default must be to block or escalate, never to wave through,” which turns an outage of the guard into an outage of the agent for the actions behind it. I accept that for writes; for a read-only path, fail open only if a wrong answer there cannot lead to a write later in the run.

How do you test a guardrail?

Test each deterministic guardrail as ordinary code with no model in the loop, and measure each probabilistic one on labeled examples.

  • Negative cases that must be refused. One per limit: the amount over the cap, the other tenant’s id, the unlisted tool name. Call the tool or the dispatcher directly.
  • A forced failure. Make the check throw or time out and assert the call is blocked, which proves the guard fails closed. Make the tool hang and assert the timeout ends the call.
  • A loop that must stop. Stub a model that requests the same tool forever; assert the budget ends the run, the trip counter ends it sooner when the tool refuses, and the output says “stopped.”
  • A wiring test. Assert the production configuration loads every guard.
  • The gate and the undo. Assert a tier 3 or 4 call is held until a person decides and that a changed argument voids the approval. Perform the action, then restore it within the window.
  • Canary fields. Chapter 17’s canary field is a schema field “constrained to be empty,” so a response that fills it fails validation.

A probabilistic guard gets the report format that 2026 evaluation used: sample sizes, a false-block rate on real traffic and a miss rate on real cases. Introduce one the way you would any risky change, as a shadow deploy followed by a canary: log what it would have blocked, read the log, then enforce for a slice of traffic, which is the canary rung.

For the stop path, I suggest a drill on staging: trip the kill switch mid-run, measure the time to halt, and check that nothing is half-applied.

What do guardrails not fix?

Guardrails don’t shrink what a permitted action can destroy, they don’t catch a wrong action that is inside every limit, and some of their costs are unmeasured. Whether the agent should hold the tool at all is a question about the blast radius of an AI agent, and no check answers it.

In-limit harm is the second. A cap of fifty permits any number of wrong refunds of forty-nine unless a quota exists, and a quota permits its own total.

Signatures are where the layer costs more than it saves. By the tiers the book adopts, every refund waits for one, and a desk that issues many a day will produce the fatigue the measurements describe. The book’s own note calls the tier boundaries “judgment calls.” A team that lets small refunds run on a cap, a daily quota and a log has made the tier depend on the amount, and should put that rule in code and write down the loss it accepted.

Three things are unmeasured as far as I could find: the latency of deterministic guards in a neutral source, approval rates across more than one product, and any independently verified cost runaway, where the public record I found is self-reports.

Three things to build first

For the ticket that says “add guardrails,” build these three first, in the order of the five places. First, keep production credentials out of the agent’s process, then a scoped credential and an undo window in the downstream system, then limits inside every write tool, each proven by a negative test. Second, budgets on the loop, a timeout on every call and a stop outside the process, proven by a stub model that loops and a drill on staging. Third, a signature gate on tier 4 actions only, with tier 3 in a review queue, triggered by code and proven by a test that the call is held and the approval binds to the exact arguments.

For every sentence in your prompt that begins with “never,” some line of code should be able to say no.

Chapter 17, “Security, Safety, and Guardrails,” develops the four families, fail-closed and the cost of the layer (in the full book); Chapter 3, The Agent Loop (in the full book) covers stop conditions and budgets. The AI agent security guide places this post beside its neighbors, or you can see the formats.

Questions readers ask

Is a rule in the system prompt a guardrail?
No. A rule in the system prompt can lower how often the model attempts the forbidden action, and nothing halts the run when the rule is broken. Keep the rule, since it is free to write, and list it as a request. The guardrail is the code that refuses the action whatever the model decided.
Are guardrail libraries or products enough?
A library that inspects input and output text covers that layer only. For an agent, the limits that matter most sit in other places: inside each tool, in the credential the agent holds, in the downstream system's own permissions and in the loop's budgets. A library can help with the first group and cannot supply the second.
Where should a guardrail be enforced: in the prompt, the model, the tool or the downstream API?
Work through five places in order. Start with what the agent's process can reach at all, then put authorization in the downstream system, argument limits and quotas inside the tool, and tool allow-lists, approval gates and budgets in the harness. Use a model or a classifier last, and never as the only check in front of an action.
What is an AI agent kill switch?
The book defines it as one obvious, fast, tested way to stop every running loop and agent at once. This post adds a placement rule: it should live outside the agent's own process, in the orchestrator, a gateway or credential revocation, so that it works when the agent's own code is the thing that is stuck.
How do you test an AI agent guardrail?
Test deterministic guardrails like any other code, with no model in the loop: call the tool with an argument that must be refused, force the check to throw and confirm the call is blocked, and stub a model that loops to confirm the budget stops it. Measure a model-based guard on labeled examples and report its false-block and miss rates.

Sources

  1. OpenAI (2025). A practical guide to building agents
  2. OpenAI (undated; read October 2026). Guardrails (Agents SDK documentation)
  3. OWASP Gen AI Security Project (2025). LLM06:2025 Excessive Agency
  4. Anthropic (2026). How we contain Claude across products
  5. Anthropic (2026). How we built Claude Code auto mode: a safer way to skip permissions
  6. Alex Wauters (2026). Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays
  7. Mrinank Sharma et al. (2025). Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
  8. Mahmoud Abdelwahab, Railway (2026). Your AI wants to nuke your database. Guardrails fix that
  9. Simon Sharwood, The Register (2025). Vibe coding service Replit deleted user's production database, faked data, told fibs galore