Home / Blog / Patterns and multi-agent systems / Add Approval Gates to AI Agents by Consequence,…

Patterns and multi-agent systems

Add Approval Gates to AI Agents by Consequence, Not Confidence

Add approval gates to AI agents by what a wrong action costs, not by model confidence. Tier each tool, design the card, test that people still read.

By Enrique Gutiérrez · Published · 17 min read

To add approval gates to AI agents, key each gate to what the action costs when it is wrong, never to how sure the model says it is. Sort every tool into a consequence tier, gate the top two (and reversible actions until they are logged, the classifier’s convention), show the approver a decision rather than a transcript, decide what happens when nobody answers, and test that people still read.

A gate that nobody reads is a log line with a delay. Most of this post is about keeping the gate from becoming one, because the measurements we have on human approvers are not kind.

The tiers come from Chapter 12, Oversight and Autonomy (in the full book), section “Approval Gates and Escalation.” The guide to AI agent guardrails places the gate among the other checks in code; this post designs the gate itself.

What is an approval gate, and where does it sit?

An approval gate is a point where the run halts until a person approves, rejects or edits the proposed action. It sits in your code, after the model has chosen a tool call and before anything executes. The gate fires on what the action is, so the model never decides whether to ask.

That placement has a name in practitioner writing. Dex Horthy’s “12-Factor Agents” asks every framework for the ability to pause and resume “ESPECIALLY between the moment of tool selection and the moment of tool invocation” (the capitals are his). The book states the rule that governs the placement: “the gate must live in your code, never in the agent’s judgment.” If the model decides at run time whether its own action needs a person, a prompt injection can talk it out of asking.

Human in the loop, in plain terms, means a person decides inside the run, before the action. Human on the loop means a person watches and can stop a run that does not ask (the glossary keeps both). A gate is the in-the-loop instrument. OWASP’s entry on excessive agency recommends the same control: “Utilise human-in-the-loop control to require a human to approve high-impact actions before they are taken.”

A gate is also different from two things it is often confused with.

An escalation trigger fires on how the run is going: repeated failure, a budget nearly spent, a suspected injection. An ask-a-human tool is one the agent may choose to call when it is unsure what you meant. Both are useful. Neither is a gate, because the agent’s state decides whether they fire.

Why gate by consequence and not by the model’s confidence?

Because consequence is a fact about the action, and stated confidence is a claim by the thing you are checking. The book’s tiers grade the action: “classify everything the agent can do into consequence tiers—read-only, reversible, externally visible, irreversible—and attach the oversight to the tier.” Its corollary is one line: “the agent’s confidence plays no part in it.”

The four consequence tiers, keyed to what an action costs when it is wrong and never to the model’s confidence.
Figure 12.2 The four consequence tiers, keyed to what an action costs when it is wrong and never to the model’s confidence. Read-only actions run freely; reversible ones run provided they are logged and can be undone; externally visible ones earn a review queue; and irreversible ones (in accent, at the top) wait for a signature, every time. Consequence rises up the stack to the one gate you cannot afford to skip. Reuse this diagram

Two arguments support the corollary.

The first is calibration. The book cites a production analysis in which a claimed confidence of ninety percent corresponded to “something closer to seventy-five percent real accuracy,” and its own footnote calls those figures “illustrative of the gap, never constants.” The second argument is stronger and needs no number: a confidence score is model output, and model output is what an attacker or an odd input can move. A gate that opens when the model feels sure can be opened by making the model feel sure.

The confidence idea has a fair version, and one top-ranking guide describes it: an agent that “runs fully automated for high-confidence outputs (say, ≥0.85 confidence on a structured schema) and routes low-confidence items to a human queue,” which the author recommends for “classification, routing, triage” (Rioja, undated). For work with no side effect, a score you have checked against your own labels is a reasonable way to decide which items a person looks at. The rule this post adopts keeps that use and draws one line: confidence may order items inside a tier; it never lowers a tier. A refund that the model is sure about is still money out the door.

How do you add approval gates to AI agents in an hour?

Sort the tools first. Ask four questions of each tool, in this order, and stop at the first answer that settles the tier. The order is the consequence tier classifier’s own choice, and so is one convention in item 4, queuing an undoable action until it is logged; the policy for each tier is the book’s.

  1. Does it change anything outside the model’s own output? If not, it is tier 1, read-only: it runs with no gate.
  2. Does it move money, delete data or touch the public? If so, it is tier 4, irreversible or high-stakes: a signature, every time.
  3. Will a third party see it, such as a customer, a vendor or another team’s system? If so, it is tier 3, externally visible: a review queue.
  4. Can the change be undone, and does the system log enough to undo and audit it? Yes and logged is tier 2, reversible: it runs, logged. Undoable in principle but not logged goes to the review queue until logging exists. Not undoable is tier 4.

Two conventions keep the sort consistent. One named outside recipient is tier 3; the public, or a send to a whole list, is tier 4, because it touches the public. And an action is graded by what it triggers downstream: a merge to a branch that nobody deploys from is tier 2, while the same merge to a branch that deploys to production touches every user and is tier 4.

The book admits the boundaries are “judgment calls.” When a team wants an exception, such as letting refunds under a cap run without a signature, the exception belongs in code, with a cap and a daily quota, and the accepted loss written down. An exception that lives only in the approver’s habits is not a rule.

What does the tiering do to an agent that gates everything?

It cuts the requests a person sees by about 95% in this worked example, and it changes which requests those are. The example, with illustrative numbers, is an AI agent for internal tools with human approval: a platform team’s ops agent with seven tools. Today the rule is “every tool call asks for approval,” so all seven are gated.

The classifier below opens on that agent. Each action carries its answers to the four questions, and “asks for approval today” is set to yes on all seven.

With JavaScript on, the Consequence tier classifier runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Consequence tier classifier on its own page to share a result by link.

Tool Illustrative requests a week Tier Gate after tiering
read_dashboards 600 1 read-only none
open_ticket 80 2 reversible, logged none; logged
scale_staging_replicas 40 2 reversible, logged none; logged
edit_runbook_page 20 2 reversible, not yet logged review queue until edits are logged
message_vendor_support 15 3 externally visible review queue
rotate_api_key 2 4 (cannot be undone) signature
drop_staging_table 1 4 (deletes data) signature

Under the current rule a person sees 758 requests a week. After tiering they see 38: 3 signatures and 35 queued items, 95.0% fewer. Once runbook edits are logged and restorable, the queue falls to 15 and the total to 18, 97.6% fewer. The tool’s headline counts the tier 3 queue only (“2 signatures, 1 queued”) and lists the unlogged runbook edit among its warnings; the other warning is that a read-only action is gated today.

rotate_api_key is the interesting row. Nobody outside sees it, and it moves no money, yet it lands in tier 4 because the old key cannot be restored once consumers have broken. Reversibility catches what visibility misses.

What should the approver see?

A decision, built by code from the exact arguments. The book lists the contents: “the proposed action in plain language, why the agent wants to take it, what it touches and whether it can be undone, and the before-and-after of anything it changes.” And it names the button teams leave out: “‘reject, with edits’ turns a dead end into a course correction, and it is often the most valuable button on the panel. Send a decision, not a transcript.”

Three properties make the card trustworthy rather than merely readable:

  • Rendered from arguments. The plain-language line and the before-and-after are generated by code from the tool call. The agent’s reason is shown, labeled as the agent’s own words, because the model may have been talked into writing it.
  • Bound to what was shown. The approval covers that action with those arguments and nothing else. One practitioner guide stores “a hash of the proposed action at interrupt time” and verifies it at execution (Digital Applied, June 2026); any changed argument voids the approval.
  • Re-checked and run once. The book’s three conditions for an approval that arrives later: it “must find a run that checkpointed cleanly, must apply to the exact action the person saw rather than a silently drifted one, and must execute exactly once.” A checkpoint and an idempotency key do that work.

Here is a card template to paste into the design doc for each gated tool. When you add approval gates to AI agents without a template, the easy build is a yes/no dialog over the raw tool call: the transcript the book warns against.

APPROVAL CARD: [tool], tier [3 | 4]

Action (rendered by code):  [verb] [object] in [system], e.g.
                            "Rotate API key billing-prod-01"
Touches:                    [records / systems / people affected, with counts]
Can it be undone?           [no | yes, by (procedure), within (window)]
Before -> after:            [field: old -> new, for every changed field]
Amount or scope:            [money, rows, recipients, total]
Agent's reason (its words): "[...]"   <- shown as the agent's claim
Why this needs you:         [tier rule that fired, e.g. "cannot be undone"]

Binds to:   hash of (tool + arguments), shown as a short code
Expires:    [time]; on expiry the action does not run
Buttons:    approve | reject | reject, with edits
After:      re-check preconditions, run once, log approver and hash

Why do approvals turn into rubber stamps?

Because approving is a vigilance task, and vigilance fails when bad requests are rare and decisions are many. One vendor reported that “users approved roughly 93% of permission prompts” in its coding agent and that “approval fatigue showed up within weeks” (Anthropic, 2026). The same vendor wrote elsewhere that when a task needs dozens of actions, “users sometimes tune them out.”

A high approval rate is not proof of failure on its own, since most requests deserve a yes. The stronger evidence is about misses. In a public browser game where players approve or deny an agent’s commands under time pressure, “The average player missed 1 in 3 threats (mean accuracy 66.3%),” and miss rates “climb back up towards the end” of a session (Wauters, August 2026). The author’s own caveat matters: about 34% of the commands shown were threats, far more than in daily work.

That caveat points the wrong way for comfort. In a laboratory baggage-screening task, Wolfe, Horowitz and Kenner found that when targets appeared on 50% of trials observers missed 7% of them; at 10% prevalence they missed 16%, and at 1% they missed 30% (Nature, 2005). This is the prevalence effect: the rarer the thing you are looking for, the more often you fail to see it. Approving agent actions is not baggage screening, so read this as an analogy, but the direction it predicts is the one the game’s caveat raises: real queues, where threats are rarer, should do worse, not better.

Hospital software shows the same pattern. A review of 17 studies found that drug safety alerts “are overridden by clinicians in 49% to 96% of cases,” while noting that “Alert overriding may often be justified” and naming “low specificity” among the conditions that produce errors (van der Sijs et al., 2006). An alert that fires on everything teaches people to dismiss it.

Tiering is the remedy that acts on prevalence directly.

In the worked example, suppose one request a week deserves a no and it is one of the consequential ones. Among 758 requests it is 0.13% of what the approver sees; among 38 it is 2.6%; among 18 it is 5.6%. The book puts the principle in one sentence: “Over-gating does more than annoy. It defeats the gate.”

What happens when nobody answers?

Nothing runs. A gated action that has no decision is treated as rejected, and the run waits or closes with its state saved. A timeout that approves by default turns the gate into a delay, and an unattended run makes that delay certain.

Tier If nobody decides The run
3 externally visible The action does not run; the draft is kept; the owner is reminded; the request expires after a window the team sets Continues with work that does not depend on the action
4 irreversible or high-stakes The action does not run; the request escalates to a named second approver; on expiry the request closes Waits at that step, checkpointed; closes with state saved if it expires
Any gated tier, unattended run The request is treated as a rejection Stops at that step and reports

Practitioner defaults for the window vary widely, which is a reason to choose one deliberately.

One guide recommends “a 7-day approval TTL for ordinary operations, 24 hours for sensitive ones” without naming its practitioners (Digital Applied, June 2026). Another advises, for a customer-visible action that times out, “Do not execute. Keep draft for review,” and for destructive or financial actions, “No unattended fallback” (FeatBit, June 2026). One builder on a public forum compressed the unattended case to a line (Hacker News, August 2026): “Unattended runs can’t ask, so needs-human becomes deny.”

Whatever the window, the approval that arrives late is re-checked against the world before it runs. A key that was rotated by someone else in the meantime, or a table that has new rows, means the approved action is no longer the action the person saw.

Can you batch approvals?

Yes, when the batch is one decision. Batch only what one approval can bind to exactly: one list, its count and its total. Never offer “approve all pending” across unrelated requests, because that button turns many decisions into one click without anyone making them.

Three cases show the line:

  • A refund batch after a billing bug. Forty refunds caused by one incident are tier 4. One signature can cover them if the card shows the 40 order ids, the count, the total and the largest single amount, and if any change to the list voids the approval.
  • A morning queue of vendor messages. Fifteen tier 3 messages to different vendors are fifteen decisions. Reviewing them in one sitting is efficient; each still gets its own approve, reject or edit.
  • A plan with a destructive step. Plan-level approval, in the book’s words, means you “sign the strategy, keep interruption rights, let the steps run.” The tier 1 and 2 steps of a signed plan run under that signature; a tier 3 step still goes to the review queue, inside the signed plan; and a tier 4 step, such as drop_staging_table, still waits for its own signature when the run reaches it.

Plan approval is the strongest batching move available. One vendor reported that showing the plan up front “shifts the user’s level of oversight from the individual step to the overall strategy, which we find tends to be where users most want to exercise judgment” (Anthropic, April 2026). How far to move a task along that range is the subject of the post on the autonomy slider for AI agents.

How do you know a gate still works?

Send it requests you know should be rejected, and count how many get through. Aviation security built it into its screening machines: an FAA plan specified that screening equipment “randomly display fictitious threats, such as knives, guns, and explosives,” and that a screener’s performance be measured by “correct responses to the display of fictitious threat images” (Abraham, SPIE Newsroom, 2008). The practice is called threat image projection.

For an agent, a seeded request is a gated action that the gate’s own code created and marked as a test. It looks like a real request, such as a vendor message containing a customer’s private data or a key rotation for the wrong service. The gate refuses to execute it whatever the answer and tells the approver afterward. Tell the team the program exists; the aim is a measurement, not a trap.

Count honestly, because seeded samples are small. Twenty seeded requests, all rejected, give a 95% Clopper–Pearson interval of 83.2% to 100% on the catch rate; ten out of ten give 69.2% to 100%; nineteen out of twenty give 75.1% to 99.9%. The eval sample size calculator computes these intervals for your own counts.

Alongside the seeded catch rate, four numbers per tier and per approver tell you whether attention is still being paid. None has a published threshold, so compare each approver with their own history:

  • Approval rate. Near 100% is normal; a sudden rise after a volume increase is not.
  • Edit rate. If “reject, with edits” is never used, either the agent is perfect or nobody reads the before-and-after.
  • Time to decide. Time yourself reading three real cards; decisions faster than that are not readings.
  • Volume and timeouts. Requests per approver per day, and how many expired. Rising expiries mean the queue has outgrown its owners.

Run this review monthly, and after every model change, since the book’s advice for the dial applies to gates too: “Re-earn the dial settings after every upgrade.”

  • Every tool has a tier, written in code next to the tool, with the question that settled it.
  • No read-only tool is gated, and every tier 2 tool has a tested undo.
  • The approval card is rendered from arguments; the agent’s reason is labeled as its claim.
  • The approval binds to a hash of tool and arguments; a changed argument voids it.
  • The approved action is re-checked before it runs and runs exactly once.
  • No gated action runs on timeout; tier 4 escalates to a named second approver.
  • No “approve all pending” button; batches bind to one list, count and total.
  • Seeded requests ran this month, and the catch rate is reported with its interval.
  • Approval rate, edit rate, time to decide, volume and expiries are reviewed per approver.
  • The gate map was re-checked after the last model or tool change.

Where are approval gates the wrong tool?

Gates are the wrong tool where the approver cannot judge the action, where the action is too frequent to read, or where the damage happens before any write. Each has a better answer elsewhere.

When the approver cannot evaluate the action. One vendor’s rule is “Match isolation strength to the user’s capacity for oversight.” A non-technical approver shown a shell command is not overseeing it, and the book says so bluntly: “a signature you do not understand is a rubber stamp, whatever tier it guards.” Use a stronger boundary that needs no approval: the same vendor reported an 84% reduction in permission prompts after moving its agent into a sandbox (Anthropic, 2026).

When volume is the problem. If a tier 3 queue is drowning its reviewers, the fix is a tool that is less dangerous (a draft instead of a send, a capped amount, a restorable delete), not a second reviewer. Approving at a volume nobody can read is what the book calls review theater; reviewing intent rather than lines is the subject of the post on human-in-the-loop review theater.

When the harm is a read. Tier 1 actions run freely, so an agent that reads private data and untrusted content is stopped only at the moment it tries to send something out. The gate on the send is necessary, and it is not sufficient; the guardrails guide covers the checks around it, including a classifier that approves tool calls in place of a person and its measured miss rate.

When the approver is another agent. In a system with a lead agent and workers, a lead that approves a worker’s tier 4 call is the model’s judgment again, one level up. The gate belongs in the harness that executes the tool, whichever agent asked. The coordination costs of multi-agent systems grow with every such hand-off.

I found no published study that measures how often seeded requests are caught in a real agent approval queue, or how approval rates change with volume across more than one product. The numbers above come from one vendor’s telemetry, a public game, a laboratory task and clinical alerts. Treat them as the reason to measure your own gate, not as a forecast of it.

The gate map in one line

To add approval gates to AI agents that stay real: gate what costs the most when it is wrong, bind the approval to exactly what the person saw, let nothing run on silence, and prove each month that a bad request gets a no. As Chapter 17 puts it: “An approval gate is a scarce resource: spend it where a considered ‘no’ is plausible.”

Chapter 12, “Oversight and Autonomy,” develops gates, escalation and the autonomy dial in full (in the full book; Chapter 17 adds the security accounting). The agent patterns guide places this post among its neighbors, or you can see the formats.

Questions readers ask

What is human in the loop AI?
Human in the loop means a person decides at a defined point inside the run, before an action happens: the agent proposes, the person approves, rejects or edits, and only then does code execute it. Human on the loop means the person watches a run that proceeds without asking and can stop it. An approval gate is the in-the-loop instrument.
Should the agent decide when to ask for approval?
No. The trigger for a gate must be code keyed to what the action is. If the model decides at run time whether its own action needs approval, anything that can persuade the model, including a prompt injection, can persuade it to skip asking. An ask-a-human tool the agent may call is useful too, but it is an addition to the gates, never a replacement.
Can I use a confidence threshold for human in the loop AI agents?
Use it to sort, never to exempt. A confidence score you have checked against your own labeled data can order items inside a review queue, or route classification work that has no side effect. It cannot move a refund, a deletion or a public post out of the signature tier, because the tier is a fact about the action.
What should happen if nobody approves in time?
The action does not run. Keep the agent's draft, notify the owner, and let a tier 3 request expire after a window your team sets. A tier 4 request escalates to a named second approver and then closes with the run's state saved. An unattended run that reaches a signature step stops there.
How many approvals are too many?
There is no published threshold. Watch the signs instead: decisions faster than anyone could read the card, an edit rate of zero, and seeded test requests getting approved. When those appear, move reversible and read-only actions out of the gate before adding reviewers.

Sources

  1. Digital Applied (2026). Human-in-the-Loop Escalation Design for AI Agents (7 June 2026)
  2. Anthropic (2026). How we contain Claude across products
  3. Anthropic (2026). Trustworthy agents in practice (9 April 2026)
  4. Alex Wauters (2026). Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays (5 August 2026)
  5. Jeremy M. Wolfe, Todd S. Horowitz, Naomi M. Kenner (2005). Rare items often missed in visual searches, Nature 435:439–440, doi:10.1038/435439a
  6. Heleen van der Sijs, Jos Aarts, Arnold Vulto, Marc Berg (2006). Overriding of drug safety alerts in computerized physician order entry, JAMIA 13(2):138–147
  7. Douglas Abraham (2008). Measuring baggage screener performance using 3D fictitious threat images, SPIE Newsroom
  8. Dex Horthy (2025). 12-Factor Agents, Factor 8: Own your control flow
  9. OWASP Gen AI Security Project (2025). LLM06:2025 Excessive Agency
  10. FeatBit (2026). Human-in-the-Loop Gates for AI Agents: Approval Without Review Fatigue (13 June 2026)
  11. Alejandro Rioja (undated). Human-in-the-Loop AI Agents: When to Build an Approval Gate (undated; read 7 October 2026)