Human in the loop AI agents are sold on one promise: a person approves the work. For an agent’s output, that promise breaks once the output passes what a person can read. My position is that the signature belongs on intent, plan, irreversible actions and evidence, and that mechanical checks should cover the lines.
This is an opinion piece, so here is the opinion without padding. A policy that says “a human reviews everything the agent produces” is false at volume, and a false policy is worse than a narrower true one. The rest of the post says where the human goes instead, how to tell a real review from a performed one, and where this argument is wrong.
The argument is the middle section of Chapter 12, Oversight and Autonomy (in the full book), “Reviewing Intent and Outcomes, Not Lines.” The table of review points, the review record and the three measurements are this post’s own extension of it.
What is review theater?
Review theater is the book’s name for “diligence performed at a volume where it can no longer be real.” The approval is recorded and the reading behind it never happened. It is the review-side form of the verification gap: the distance between what was checked and what was claimed.
The book describes the scene exactly. A change arrives, four hundred lines across nine files with green tests, and the reviewer starts at line one. By the third file they are skimming. The approval, the chapter says, “now certifies mostly that you were present.”
Nobody in that scene is lazy. The agent can produce the next change in minutes, and a person’s careful reading speed does not move. Engineers describe the result in their own words. One wrote in May 2026: “We’re seeing much higher PR volume from AI-generated code, but review time hasn’t improved, so PR review has become the main bottleneck instead of coding” (Hacker News). Another, the same month, described layers of automated review that exist “all because there’s simply too much volume of code to review” (Hacker News).
A vendor’s telemetry points the same way. Faros AI reported in July 2025, from “over 10,000 developers across 1,255 teams,” that “Developers on teams with high AI adoption complete 21% more tasks and merge 98% more pull requests, but PR review time increases 91%” (Faros AI, 2025). That is one company’s customers and a correlation, and it concerns coding assistants. Read it as a reason to measure your own queue.
Why can’t the reviewer just read faster?
Because the reading rate at which review finds defects has a ceiling, and agent output passes it quickly. The only published bounds I could fetch predate agents and come from a vendor, so treat them as a starting placeholder and time your own team.
SmartBear’s summary of its study of one Cisco Systems team says “developers should review no more than 200 to 400 lines of code (LOC) at a time.” The same page reports, as SmartBear’s own research, that there is “a significant drop in defect density at rates faster than 500 LOC per hour,” and says more generally that “performance starts dropping off after about 60 minutes” (SmartBear, undated). All of it concerns human-written code, and the first figure is one team.
Here is a worked example with illustrative numbers. An agent returns ten changes of 400 lines in a day: 4,000 lines. A reviewer has two hours for review and a careful rate of 300 lines an hour, the same placeholder the team workflow post uses.
| Quantity (illustrative) | Value |
|---|---|
| Lines produced in the day | 4,000 |
| Hours to read all of it at 300 lines an hour | 13.3 |
| Lines readable in the two-hour budget | 600 |
| Share of the day’s output actually read | 15% |
| Reading rate implied if all ten are approved in two hours | 2,000 lines an hour |
Two thousand lines an hour is four times the 500 that SmartBear’s page names as the point where defect finding drops. An approval at that rate is a signature on the 85% nobody read.
The usual replies are to add reviewers or to cap the agent. Adding reviewers scales linearly against a producer that does not tire: at these placeholder numbers, reading everything needs 13.3 reviewer-hours a day for one agent. Capping the agent at 600 lines a day is honest, and for some teams it is the right answer. I come back to it under limits.
Where do human in the loop AI agents need the human?
A person’s judgment belongs at four points: the intent before the run, the plan before execution, each irreversible or high-stakes action during the run, and the evidence after it. At each of those, the thing to read is short and written for a person, and the decision changes what happens next.
The book’s version is two questions: “Did the agent set out to do the right thing, inside the scope we agreed on? And can it prove that what it did works?” The four signatures are where those two questions get asked.
Intent. Success criteria stated before the run, so that, in the book’s words, “‘done’ was defined by you rather than claimed by the agent.” The post on spec driven development with AI agents covers how to write them.
Plan. The book: “Approving five proposed steps takes two minutes and catches the misunderstood requirement while it is still a sentence you can edit.” And: “Plans, unlike diffs, arrive at human reading speed.” One vendor that moved its product from per-action prompts to plan approval wrote that this “shifts the user’s level of oversight from the individual step to the overall strategy, which we find tends to be where users most want to exercise judgment” (Anthropic, 2026). That is a vendor describing its own product; the pattern is general.
Irreversible actions. These stop at an approval gate in code. Which actions qualify, what the approver sees and what happens on silence are the subject of the post on approval gates by consequence, and I won’t re-derive them here.
Evidence. Test output, a reproduction, logs, a screenshot: in the book’s phrase, “proof a stranger could check without trusting the claim.”
Which review points are there, and what covers the rest?
The table lists nine review points, what the person decides at each, and what covers the part the person does not read. The last column says when the point occurs, and the buttons filter on it.
| Review point | What the human decides | Why a person and not a check | What covers the rest | When |
|---|---|---|---|---|
| Intent: the task and its acceptance criteria | Whether this is the right thing to build, and what “done” means | No check knows what you wanted | Nothing; this is the root | before the run |
| Declared scope | Which directories, features or records the run may touch | Scope is a choice about risk | A mechanical diff of touched paths against the declaration | before the run |
| Plan | Whether the approach is one you would accept | A wrong approach passes every test written for it | Interruption rights once execution starts | before the run |
| Irreversible or high-stakes action | Approve, reject or edit this exact action | The cost of a wrong yes cannot be undone | The gate in code; lower tiers run without a person | during the run |
| Escalation or a question from the agent | A preference or a policy call the agent cannot look up | Only you know what you meant | Triggers in the harness: repeated failure, budget, suspected injection | during the run |
| Evidence against the acceptance criteria | Whether each criterion has proof you could re-run | Someone must decide the proof answers the question asked | The test suite, type checks, linters, an independent judge model | after the run |
| Any change to a test or a check | Whether the check still tests what it did | A weakened check makes everything after it pass | Read line by line, every time | after the run |
| Code you must own | Whether you understand it well enough to change it next | Your next decisions depend on it | Read in full, or write it yourself from the agent’s guide | after the run |
| Random sample of unread changes | Whether the checks are catching what a reader would | Checks cannot audit themselves | A full line read of the sample, by someone other than the approver | across runs |
When do you still read the lines?
Read lines in four cases: a flag, a change to a test or check, code you must own, and a random sample. Everything outside those four is covered by mechanical checks and written down as not read.
A flag is a touch outside the declared scope, a criterion with no evidence, or a walkthrough that disagrees with the diff. The book makes the first one mechanical: “anything touched outside it is a flag, regardless of how good it looks.”
A test or check change is read every time, because it is the cheapest way for a run to go green. One engineer described a former employer’s firmware project in September 2026: “tests were being passed, not because the code was sound, but because the tests were altered to match the results” (Hacker News). That is one person’s unverifiable account, and it is the failure this rule exists for. The team workflow post applies the same rule and adds new entry points to it.
Code you must own is the part of the system your next ten decisions depend on, plus anything on a security boundary. The book’s instruction for it is to “read the walkthrough slowly, or reach for the inversion and write the thing yourself,” an option it calls the “Wrong default for bulk chores; strong option for the code you must own.”
The random sample audits the checks. It is the only one of the four that needs no trigger, which is why it finds what the other three miss.
What does the same two hours buy under this rule?
In the worked example, the same two hours cover all ten changes at the level of intent and evidence, with forty minutes left for lines. The numbers are illustrative.
Suppose an intent-and-evidence pass takes eight minutes a change: three for the plan and criteria, two for the scope diff, three for the evidence. Ten changes take 80 minutes. The remaining 40 minutes read 200 lines at 300 an hour, half of one change.
So the rule has a ceiling too. At eight minutes each, two hours cover at most 15 changes with no line reading at all. One full read of a 400-line change takes 80 minutes and leaves 40, which is five intent passes: its own and four others. A reviewer who owes a full read today can honestly sign five changes in all, not ten.
How do you detect review theater on your own team?
Measure three things: the reading rate each approval implies, whether a seeded defect is caught, and what a full read of a random sample finds. None of them needs new tooling, since the first comes from timestamps a review system already keeps.
Implied reading rate. Divide the lines in an approved change by the minutes between opening it and approving it. Compare that with the rate you get by timing two or three careful reviews on your own code. An approval far above your careful rate was not a line review, whatever the policy calls it. Open time overstates reading, so this measure flatters the reviewer.
Seeded defects. Plant a change that breaks a stated acceptance criterion or touches a file outside its scope, and see whether it gets a no. The gates post works out what a catch rate from a small number of seeds does and does not prove.
The sample. Have someone other than the approver read a random sample of unread changes in full. A sample is only as good as its size. If 20% of unread changes held a problem a reader would catch, a sample of five would contain at least one such change 67% of the time, and a sample of ten 89%. At a 5% share those figures fall to 23% and 40%. Containing a bad change is not the same as catching it.
The checklist below turns the same idea into signs. Tick each one that is true of your team’s last ten approvals of agent work. Any tick marks a place where the signature covers more than the reading.
- An approval’s implied reading rate was several times the team’s own careful rate.
- No acceptance criteria existed before the run, so “done” was whatever the agent reported.
- No scope was declared, so nothing could flag a file touched outside it.
- The plan was never shown to a person before execution began.
- A test or check changed in the same change and no person read that file.
- The evidence was the agent’s own summary, with no output a reviewer could re-run.
- The approver could not say what the change was for without reopening it.
- The approval record does not say what was read and what was not.
- No seeded defect and no random sample has been run since the last model or tool change.
- A change drawn for the random sample was read by its own approver.
What should an approval record say?
An approval record should say what the signature covered: which of the four signatures were given, which of the four line-reading cases applied, and what nobody read. The book’s ledger is the trust ledger kept on the human’s side, because “You are maintaining the ledger for a worker who cannot keep one.” A per-change record is this post’s extension of that idea.
The point of the record is that it makes “not read” a legitimate entry. A team that can write “lines read: none; covered by checks” stops needing to pretend, and the sample then has a population to draw from.
REVIEW RECORD: [change or run id], [date]
1. INTENT acceptance criteria written before the run? [yes / no]
declared scope: [paths, features or records / none declared]
approved by: [name] where: [link]
2. PLAN shown before execution? [yes / no / trivial task, none needed]
approved by: [name] edits made: [none / summary]
3. ACTIONS irreversible or high-stakes actions in this run: [none / list]
each approved at its gate by: [name(s)]
4. EVIDENCE criterion -> proof a reviewer could re-run:
[criterion 1] -> [test output / reproduction / log / screenshot]
[criterion 2] -> [...]
criteria with no proof: [none / list] checked by: [name]
LINES READ
flag raised? [none / out-of-scope paths / criterion with no evidence / walkthrough disagrees with diff]
read by: [name] files: [list]
test or check changed? [no / yes] read line by line by: [name]
code I must own? [no / yes] read in full by: [name]
in the random sample? [no / yes] read by: [name, not the approver]
NOT READ BY A PERSON
files or line count: [list or number]
covered by: [test suite / type checks / linters / judge model / nothing]
SIGNATURE
I approved items [1-4 as applicable]. I read [list]. I did not read [list].
[name]
Does the rule hold outside a code diff?
It holds wherever an agent produces more than a person can read, with the tiers deciding what still waits for a signature. Three cases the post has not used so far show how the table, the four cases and the record line up.
A dependency bump across 60 files. Intent is one sentence and the plan is trivial, so the record says “none needed.” Scope is the manifest and the call sites; the path diff checks it. Evidence is the build and the test suite. No lines are read unless a flag fires or a test file changed, and the change goes into the pool for the sample.
A migration that drops a column. Running it is an irreversible action, so it waits at a gate. The script is also code you must own, so it is read in full; a 60-line script takes about 12 minutes at the placeholder rate. Evidence is a dry run on a copy with row counts. This case collects all four signatures and one line-reading case.
An agent that drafts 200 support replies a day. Sending is externally visible, so the review queue from the gates post applies. The capacity test applies to that queue as well: 200 replies in a two-hour budget is 36 seconds each. If that is not a real reading, the choices are the ones above. Sign the policy the replies are written against, sample, and move the task along the autonomy dial only on evidence. Where a task sits on the autonomy slider for AI agents is a separate decision from this one.
If checks keep improving, why keep a human at all?
Because checking and steering are different jobs, and only checking is being automated. This is the strongest objection to the whole post, and it deserves its own words.
One commenter put it in September 2026: “Humans have been shipping systems that no one person understands for a long time,” and “AI agents will be reviewing code soon enough” (Hacker News). The first half is true. Nobody reads the compiler’s output either.
The book’s answer borrows a distinction from a talk by Geoffrey Litt. Understanding in order to verify is a thumbs-up question, and tooling keeps taking those. Understanding in order to participate is what lets you choose the next task. Simon Willison’s note on the talk summarizes it: “you need to understand the code to a depth that enables you to participate further with the model” (Willison, July 2026). The book closes the section on it: “Verification you can delegate, increasingly to the agent itself. Participation is the part you can only do, or lose.”
So I accept the objection for the checking role and reject it for the signatures on intent and plan. A reviewer model is welcome as one more check. Two conditions apply. It should run in a separate context from the worker, since the book calls an agent judging its own run “the weakest verifier you can buy.” And an LLM-as-a-judge needs its agreement with human labels measured before its verdicts count.
What does this argument cost, and where is it wrong?
It costs the side benefits of line review, it leans on criteria and checks that may be weak, and it does not apply where volume is low or the code is the kind you must own. I found no published comparison of defect escape rates between line review and intent-plus-evidence review of agent-written work, so the position rests on capacity arithmetic and on the book’s reasoning, not on a measured outcome.
It moves the theater upstream if the criteria are vague. An approved spec that says “make checkout faster” gives the evidence nothing to be checked against. Then the four signatures are four rubber stamps. The book’s line applies at every level: “a signature you do not understand is a rubber stamp, whatever tier it guards.”
The walkthrough is written by the party under review. It is a reading aid and a claim. Only output a reviewer could re-run counts as evidence.
Line review did more than find defects. It spread knowledge of the codebase and taught newer engineers. Dropping it for most changes drops that too, and the “code you must own” case only partly replaces it.
Low volume needs none of this. If your careful rate covers what the agent produces, read it. Capping the agent’s output to what reviewers can read is a legitimate choice, and for a small team on a critical system it may be the best one.
Some rules name the failure and still require oversight. The EU’s AI Act asks that people overseeing a high-risk system be enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)” (Regulation (EU) 2024/1689, Article 14). That provision concerns high-risk systems under that regulation. Whether a given review process satisfies any rule is a question for whoever advises you on it.
Checks written by agents can share the worker’s blind spot. A second model reading the first model’s work may miss the same thing for the same reason. The comparison of multi-agent systems books finds the same warning in the classic theory: agents that share a model behave in correlated ways.
The signature, restated
For human in the loop AI agents to mean something, the approval has to say what the person read. Sign the intent, the plan, the irreversible actions and the evidence; read lines on a flag, on a test change, in code you must own and in a random sample; and write down the rest as not read. A narrower claim that is true beats “reviewed” on a diff nobody could have finished.
To place the gates, the consequence tier classifier sorts actions by what they cost when wrong, and the agent verifiability scorecard asks the wider question of how you would know a run worked.
Chapter 12, “Oversight and Autonomy,” develops gates, intent review and the dial in full (in the full book). The agent patterns guide places this post among its neighbors, or you can see the formats.
Questions readers ask
- What is human in the loop AI?
- Human in the loop means a person decides at a defined point inside the run, before an action happens: the agent proposes, the person approves, rejects or edits, and only then does code execute it. Human on the loop means the person watches a run that proceeds without asking and can stop it. Reviewing an agent's finished output is a third thing, and it is the one that fails first as volume grows.
- What is review theater?
- Review theater is the book's term (Chapter 12) for diligence performed at a volume where it can no longer be real. The approval exists, the reading behind it does not. The usual form is a large agent-written change approved in less time than its lines could have been read, on the strength of green checks that say nothing about whether the right thing was built.
- Should a human still read every line an agent writes?
- Yes, while the volume fits. If a reviewer's careful reading rate covers everything the agent produces, line review is honest and nothing here asks a team to stop. Once output passes that rate, read lines in four cases only (a flag, a test or check change, code the reviewer must own, a random sample) and record the rest as not read.
- Can a second AI agent do the review instead?
- A second model can do part of the checking, and it should be a different context from the one that did the work. It cannot own the decision about what to build, and its own accuracy has to be measured against human labels before its verdicts count. Treat it as one more mechanical check under the four signatures, never as the signature.
- Does reviewing intent meet a requirement for human oversight?
- That depends on the rule, and nothing here is legal advice. What can be said in general is that an approval record stating what the person read, what the checks covered and what nobody read is easier to defend than an approval that implies a full reading which did not happen.
Sources
- SmartBear (undated). Best Practices for Peer Code Review (undated; read 7 October 2026)
- Faros AI (2025). The AI Productivity Paradox Research Report (23 July 2025)
- Anthropic (2026). Trustworthy agents in practice (9 April 2026)
- Simon Willison (2026). Understand to participate (2 July 2026)
- European Parliament and Council (2024). Regulation (EU) 2024/1689, Article 14: Human oversight