What is a human in the loop approval policy?
A human in the loop approval policy is a written rule that says which of an agent’s actions run on their own and which wait for a person. The book’s version keys every rule to what an action costs when it goes wrong: read-only actions run, reversible ones run with a log, externally visible ones queue for review, and irreversible ones wait for a signature.
This tool builds that policy for your agent. You list the actions it can take, answer up to four questions about each, and get a tier, a gate and the reason for both. The method comes from Chapter 12 of the book, Oversight and Autonomy, in the section “Approval Gates and Escalation.” The chapter opens with a picture worth keeping: the signing limit, the policy in most companies that says who may sign for what, “calibrated to consequence, and revised as trust accumulates.”
How do you sort an agent’s actions into consequence tiers?
You sort them by asking what each action costs when it is wrong, never how good the model is. The book’s instruction is to “classify everything the agent can do into consequence tiers—read-only, reversible, externally visible, irreversible—and attach the oversight to the tier.”
The chapter starts with a sorting exercise, and the tool loads it as its default example. A support agent has five tools: look up an order, draft a reply, send the reply, issue a refund, and delete a customer record. Most readers sort them correctly without thinking, and the book points at what they consulted: “You did not ask how capable the model is, or how well it has behaved this month. You asked what each action costs when it is wrong.”
Each tier earns a different kind of oversight, and the book states all four in four sentences. “Read-only actions run without ceremony; gating them buys no safety and teaches the person approving them to stop reading. Reversible actions run freely too, provided the system logs enough to undo and audit them. Externally visible actions, anything a third party will see, deserve a review queue. Irreversible or high-stakes actions wait for a signature, every time.”
| Tier | What it means | Example from the book | Gate the tool generates |
|---|---|---|---|
| 1 Read-only | Changes nothing outside the model’s own output | Look up an order | Run; no gate |
| 2 Reversible | Can be undone, and the system logs enough to undo and audit it | Draft a reply | Run; log for undo and audit |
| 3 Externally visible | A third party will see it | Send the reply | Review queue, asynchronous |
| 4 Irreversible or high-stakes | Money, deletion, the public, or anything that cannot be undone | Issue a refund; delete a record | Signature, every time; synchronous block |
Which questions does the tool ask about each action?
The tool asks four questions per action, in an order that settles the clear cases first. Does the action change anything outside the model’s own output? If not, it is tier 1 and the other questions disappear. Does it move money, delete data, or touch the public? If so, it is tier 4, whatever else is true.
Will a third party see it? If yes, it is tier 3. Only then does the tool ask whether the change can be undone. A change nobody outside sees but nobody can undo is tier 4; one that can be undone and is logged is tier 2. A fifth question, whether the action asks for approval today, does not affect the tier. It lets the tool compare your current gates with the ones the book would set and flag the difference.
Three choices here are the tool’s, not the book’s, and the page labels them. Visibility is checked before reversibility, so a sent reply lands in tier 3 as it does in the chapter’s exercise. An action that could be undone in principle but is not logged is queued until the logging exists, because the book grants reversible actions their freedom only “provided the system logs enough to undo and audit them.” An action you have not finished answering is treated as tier 4 until you do. Failing closed costs you a warning; failing open could cost you a refund.
Why does the model’s confidence play no part?
The model’s confidence plays no part because the self-report is untrustworthy in the worst direction: a model that says it is sure is often less accurate than it claims, while what an action costs when it is wrong is a fact you can check before anything runs. The book calls this a corollary “that sounds obvious and is violated constantly: the agent’s confidence plays no part in it.” Letting an agent act freely when it says it is sure feels natural, because that is roughly how you manage people.
The chapter cites one production analysis in which a claimed confidence of 90 percent corresponded to something closer to 75 percent real accuracy, and three such steps chained together delivered roughly 42 percent end to end. Those figures are illustrative of the gap, not constants. The durable point is the one the book draws from them: “So the rule stands on consequence, which has the great virtue of being a fact about the world rather than a feeling about the model. The refund needs a signature because it is money. How sure the agent felt does not enter into it.”
The tool makes this visible in two ways. No question asks about confidence or capability. And if you paste your current approval rule into the optional field and it keys the gate to the model’s confidence, certainty or being sure, the result flags it. Everyday phrases such as “make sure” or “the probability of fraud” do not count. The compounding error calculator shows why a small calibration gap per step becomes a large one over a run.
Where should the approval gate live?
The approval gate should live in your code, in the seam between the model proposing an action and your program executing it. The book places it “after the model has chosen, before anything runs.” Dex Horthy’s 12-Factor Agents makes the same point from the builder’s side: you need to interrupt an agent “ESPECIALLY between the moment of tool selection and the moment of tool invocation.”
One rule outranks the rest, and the book sets it in bold: “the gate must live in your code, never in the agent’s judgment.” If the model decides at runtime whether its own action needs approval, then anything that can persuade the model can persuade it to skip the asking, and prompt injection is exactly that kind of persuasion. The practitioner guide the chapter draws on puts it in one line the book quotes: “The gate fires based on what the action is, not on what the model inferred about the request.”
That is why the tool’s output is keyed to action names. Each line of the policy can become a lookup table in your executor: tool name in, gate out, with no model in the decision. An approval gate built this way is also cheap to audit, since the whole policy fits on one screen.
What triggers an escalation, and what does the person receive?
Escalation triggers fire on how a run is going, where gates fire on what an action is. The book’s short list: “Repeated failure on the same step; an input outside the policies the agent was briefed on; signals of a possible injection; a budget (time, money, steps) close to spent; and, if you like, a dip in the agent’s self-reported confidence, useful as one signal among several and never as the only one.”
One asymmetry governs what happens next. “A low-stakes question can sit in a queue for asynchronous review,” while “an irreversible action, or a suspected injection, blocks synchronously until a person clears it.” The tool writes this into the tier 3 policy, which queues asynchronously but blocks if injection is suspected. The chapter adds a warning to budget for: “a gate without durable state is a gate that works only while you happen to be watching.” An approval that arrives the next day must find a run that checkpointed cleanly.
What the person receives decides whether the gate works. The tool generates a handoff template for every gated action: the proposed action in plain language, why the agent wants to take it, what it touches and whether it can be undone, and the before-and-after of anything it changes. It also adds a third button. The book calls “reject, with edits” often the most valuable one on the panel, and its rule for the whole package is short: “Send a decision, not a transcript.”
How does the autonomy dial change the policy?
The autonomy dial changes how much the agent may do between moments of your attention, and the tool adjusts the lower tiers to match. The book names four positions: every consequential action signed; standing permissions, where action types are pre-approved; plan-level approval; and monitored autonomy against budgets with a kill switch within reach.
At the tightest position the tool upgrades tier 3 from a review queue to a signature. Tiers 1 and 2 stay ungated even there, because the tool reads “consequential” as externally visible or irreversible, and a logged change you can undo is neither; that reading is the tool’s, not the book’s. With plan-level approval, tier 3 still queues, inside the plan you signed. With monitored autonomy, tiers 1 and 2 are marked for review after the fact. Tier 4 does not move at any position; the chapter’s last line is “What has not changed, at any position of the dial, is whose signature it is.” If you choose monitored autonomy while tier 4 actions exist, the tool reminds you that Chapter 13 prices that position in full.
The dial moves on evidence, and the book adds a caution specific to agents. A model upgrade can raise or lower competence overnight, so the evidence for a setting expires with the model version it measured. “Re-earn the dial settings after every upgrade.” The copied Markdown policy carries that line so the reminder travels with the document.
What does a worked example look like?
Load the book’s five support-agent tools and the classifier produces the policy the chapter describes: looking up an order runs with no gate, drafting a reply runs with a log, sending it goes to a review queue, and issuing a refund or deleting a record waits for a signature. look_up_order reads only: tier 1, no gate. draft_reply changes something that can be thrown away and is logged: tier 2, run with a log. send_reply reaches a customer’s inbox: tier 3, review queue. issue_refund moves money and delete_record deletes data: both tier 4, a signature every time.
Now change one answer. Mark look_up_order as asking for approval today, and the tool warns that gating read-only actions “buys no safety and teaches the person approving them to stop reading.” Mark issue_refund as not gated, and the tool flags a tier 4 action running without a signature. Paste a rule such as “ask a human when the agent is less than 80 percent sure,” and the tool flags the confidence clause. Every warning quotes the chapter, so the policy you share carries its own justification.
Where does this method stop working?
The method stops working at the boundaries between tiers, which are judgment calls that depend on your context, and it leaves two things uncovered, the broad middle of ordinary reversible changes and security exposure, so treat the tiers as a starting policy rather than a complete defense. The book says so in a footnote: “The tier boundaries are judgment calls; the ordering principle—reversibility and visibility, never model confidence—is the durable part.” An internal message to a colleague is visible to a third party in one company and routine in another. The tool asks; you answer for your context.
It also leaves the broad middle uncovered. Hundreds of ordinary, reversible changes pass no gate, and no person can read them line by line. The book’s answer for that middle is a different level of review, covered in the next section of the chapter and in the article on human in the loop AI agents and review theater. Over-gating is the opposite failure: “Over-gating does more than annoy. It defeats the gate.” See review theater for what happens when approvals become a rhythm.
Finally, tiers describe consequence, not exposure. A read-only tool that reads private data next to a tool that can send messages is a security problem the tier system does not see; the lethal trifecta audit checks for it. For the wider argument, read how to add approval gates to AI agents by consequence, not confidence, the autonomy slider for AI agents, AI agent guardrails and the agent patterns guide. If you are still deciding whether the task needs an agent at all, start with should this be an agent? The full method is in Chapter 12, Oversight and Autonomy.
Questions readers ask
- Why can’t the agent decide for itself when to ask for approval?
- Because anything that can persuade the model can then persuade it to skip the asking: an oddly worded task, a poisoned document, a prompt injection. The book’s rule is that the gate must live in your code, never in the agent’s judgment, and that it fires on what the action is.
- Why does the book say gating read-only actions makes things less safe?
- Every approval spends human attention. Gate actions that cannot hurt anyone and the approvals become a rhythm; the person stops reading them, and the one request that deserved a hard look gets the same reflexive click as the forty before it.
- What should the human actually receive when a gate fires?
- A decision, not a transcript: the proposed action in plain language, why the agent wants to take it, what it touches and whether it can be undone, and the before-and-after of anything it changes. Offer more than yes or no; “reject, with edits” is often the most valuable button.
- Can I use the model’s confidence score as one of the escalation triggers?
- As one signal among several, yes, but never as the only one and never as the thing that decides the tier. Stated confidence is poorly calibrated; the book cites an analysis where a claimed 90 percent meant about 75 percent real accuracy.
- Does moving to monitored autonomy remove the signatures?
- No. The dial changes how much the agent may do between moments of your attention, mostly in the lower tiers. Irreversible and high-stakes actions keep their signature at every position.
Sources
- Digital Applied (2026). Human-in-the-Loop Escalation Design for AI Agents
- Dex Horthy (HumanLayer) (2025). 12-Factor Agents, Factor 8: Own your control flow
- Anthropic (2026). Trustworthy agents in practice
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents