The autonomy slider AI agents sit on sets how much an agent may do between moments of a person’s attention. The book gives it four positions, tightest first: every consequential action signed, standing permissions, plan-level approval, and monitored autonomy. Each task gets its own position, and a position changes on evidence.
A note on names before anything else. In the book the four positions belong to the autonomy dial, the engineer’s name, built in Chapter 12, Oversight and Autonomy (in the full book). “Autonomy slider” is the product’s name for the same instrument, adopted in Chapter 25, Transversal Recipes: Big Patterns That Cut Across Industries (in the full book), whose footnote says the term “is credited to Andrej Karpathy’s talk ‘Software Is Changing (Again)’ (YC AI Startup School, 2025)”.
After this post you can place each task your agent does at one of four positions, write down the measured evidence that moves it one notch (and what moves it back), and sketch the control your user sees there.
Both the positions and the rule that evidence moves them are the book’s. The promotion table, the product-side table and the per-task record are this post’s additions, and each one is marked where it appears.
What are the four positions?
The four positions are every consequential action signed, standing permissions, plan-level approval, and monitored autonomy. The glossary lists them in one line: “Four working positions: every consequential action signed; standing permissions with escalation at the envelope’s edge; plan-level approval; monitored autonomy against budgets.”
In the table, the first two columns after the name restate Chapter 12. “When it fits” quotes the chapter where it speaks. The last column, how each position fails, is my reading of the chapter and of the sources later in this post.
| Position (the book’s name) | Who decides what | What the human sees | When it fits | How it fails (this post’s reading) |
|---|---|---|---|---|
| 1. Every consequential action signed | The agent proposes; a person signs each consequential action. Read-only actions and logged reversible ones run without a signature | One decision card per consequential action | “right for new agents, new task classes, and anything near money, deletion, or the public” | Signatures become a rhythm and stop being read |
| 2. Standing permissions | A person pre-approves action types (always allow, always ask, never). The agent “runs freely inside the envelope, escalating at its edge” | The escalations at the envelope’s edge, the always-ask types, and a log of what ran | A class of action that, in the chapter’s words, “has been boring for long enough” | The envelope is drawn by action name and misses the effect; an allowed action is still the wrong one |
| 3. Plan-level approval | A person signs the plan: “sign the strategy, keep interruption rights, let the steps run” | The plan before the run, progress and a way to interrupt during it, the result with its evidence after | Multi-step work where the plan is short enough to read | The plan is approved unread, or the run drifts from the plan it showed |
| 4. Monitored autonomy | “the agent runs unattended against budgets, its work reviewed after the fact” | Alerts, budget stops, a sample of finished work, a kill switch | Only with budgets, monitoring and “a kill switch within reach” | Nobody reads the sample, and the budget becomes the only brake |
The first box in the book’s figure is labeled “approve every action,” while the chapter’s prose and the glossary say every consequential action. This post follows the prose, because the chapter’s own tier rule lets read-only actions run “without ceremony”.
One rule holds at every position. The chapter says irreversible or high-stakes actions “wait for a signature, every time.” Moving a task looser changes how the ordinary steps are overseen, and the signature on moving money or deleting a record stays where it was.
A word on direction. The book’s figure stacks the positions with the tightest on top, so its caption has a task that “moves down the dial only on evidence”, and down there means looser. This post’s tables use the opposite convention: up means looser and down means tighter.
Why does the position belong to the task?
The position belongs to the task because the cost of a wrong action is a property of the task, and one agent does tasks of very different cost. The chapter defines the dial as “how much an agent may do between moments of your attention, set per task rather than built into the system.”
The book presses the point with a working day: it describes one practitioner who codes the core of his work by hand and lets an agent run for hours on secondary chores. Its conclusion: “it is set per task, and the same person on the same day should be running settings from both ends.”
Behind that passage is Geoffrey Litt’s essay “Code like a surgeon” (October 2025). Litt writes that his two work patterns remind him of “Andrej Karpathy’s ‘autonomy slider’ concept,” and warns: “It’s dangerous to conflate different parts of the autonomy spectrum”.
Chapter 25 says the same of whole products. Its digital coworker recipe sets autonomy “per action,” and gives the example in one clause: “the password reset runs free while the account closure waits for a signature, inside one and the same coworker.”
So a single “autonomous mode” switch for the whole product is the wrong unit: a product owner needs one setting per task. Which of those decisions belong to the product owner is the larger subject behind AI agents for product managers.
Chapter 25 places its eight recipes along the slider as starting points, from recipes that act on real systems behind gates to one that runs unattended in a sandbox. Its prose describes the personal knowledge base as an assisted tool, with ingestion “one source at a time, with its owner staying involved.” The post on AI agent use cases for startups walks the recipes one by one, so I will leave them there.
Is the choice really between babysitting and full auto?
The choice has four options, and public discussion usually names only the two at the ends. In the forum threads I read for this post, builders argue for one extreme or the other, and the two middle positions go unnamed.
Here is the tight end, rejected. One builder wrote in March 2026: “You can’t have an agent that asks permission for every action, you’d just be babysitting it all day.” A reply under it gave the opposite reflex in five words: “Don’t give it root access” (Hacker News).
Here is the loose end, embraced. A commenter in July 2026, spelling as posted: “Using an agent without YOLO mode is not wort it. The way I rather do it is tightly control the output by skills written yourself, prompts, plans, etc.” (Hacker News). The second sentence keeps some control back, through plans and instructions written in advance.
A third commenter asked for the middle outright, in June 2026: “I would want it to be safe enough to run 80 to 90% of typical development work without supervision, and then have an escape hatch that allows doing other things with human supervision” (Hacker News).
The reason the middle gets skipped shows up in a fourth comment, from September 2026: “You don’t want to approve 100 windows to get a task done” (Hacker News). When the only tight setting on offer is a prompt per action, people leave it for the only other setting on offer.
These quotes come from users and builders of coding agents, so treat the pattern as a report from one corner of the field.
What does the evidence say about approving every action?
The evidence says per-action approval degrades as volume rises, and that people who gain experience change the form of their oversight. Most of it comes from one vendor’s data about its own coding product, so it describes that product’s users.
Parasuraman and Manzey’s 2010 review in Human Factors found that “Automation complacency occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator’s attention.” They add that it “is found in both naive and expert participants and cannot be overcome with simple practice” (Parasuraman and Manzey, 2010). Their review predates language models.
Anthropic’s “Measuring AI agent autonomy in practice,” dated February 18, 2026, analyzed interactive sessions of its own coding agent for users who signed up after September 19, 2025. Its finding on tenure: newer users, with fewer than 50 sessions, “employ full auto-approve roughly 20% of the time; by 750 sessions, this increases to over 40% of sessions” (Anthropic, 2026).
The same study counts interruptions, and the unit changes from sessions to turns. Users with around 10 sessions interrupt the agent in 5% of turns; more experienced users interrupt in around 9% of turns. So in that product, experienced users approve less often in advance and step in more often while the agent works.
The study states its own limits, and they belong beside the numbers. The product’s “default settings require users to manually approve each action, so part of this transition may reflect users configuring the product to match their preferences”. The authors “can only analyze traffic from a single model provider,” and the coding data covers “a single product that is overwhelmingly used for software engineering.”
A separate page from the same vendor, published May 25, 2026, reports what happened at the tight end: “users approved roughly 93% of permission prompts” (Anthropic, 2026). That number counts prompts, a third unit. The post on approval gates by consequence covers that measurement and the research on missed rare events, along with which actions deserve a gate at all.
The chapter’s sentence is “Over-gating does more than annoy. It defeats the gate.” A signature given at a volume nobody can read is what the book calls review theater.
For this post, the consequence is narrow: position 1 has a capacity, and a task that exceeds it has already left position 1 in everything but name. The argument for signing intent and evidence in place of every line is made in the post on human-in-the-loop review theater.
What evidence moves a task one notch?
The book names the kinds of evidence and gives no thresholds. It lists “runs completed cleanly, escalations that were sensible, evaluation scores on the task class,” and “the measured error rate on a held-out set once you have one”.
“Start every new task class tight.” Then: “When a class of action has been boring for long enough, grant it standing permission and spend your recovered attention on the tiers that still bite.”
On the other direction the book is brief. The chapter’s opening says a position should move “deliberately, on evidence, in both directions.” Its only stated trigger for tightening is a model upgrade, which “can raise or lower competence overnight, in different directions on different tasks”. The instruction is one line: “Re-earn the dial settings after every upgrade.”
A forum reply states the underlying idea well. Under a May 2026 comment carrying the claim that automation beyond a human in the loop is “just theater,” another commenter wrote: “Like all stochastic processes, LLM errors can be quantified. That makes each use case a risk-reward tradeoff” (Hacker News).
Everything in the table below is this post’s own. It turns the book’s kinds of evidence into criteria per notch, adds tightening triggers beyond the model upgrade, and gives each criterion a number. Every number is an illustrative starting value. The rule of one notch at a time is mine as well.
The last column names the move, and the buttons filter on it.
| What you must have measured | Where it comes from | Illustrative starting value | The move |
|---|---|---|---|
| Signed actions of this type on the current model version | Approval records | At least 200 | up to standing permissions |
| Share of those edited or rejected | Approval records | At most 2%, and no rejection was for harm | up to standing permissions |
| Share reversed or corrected after running | Undo log, corrections, support tickets | At most 1% | up to standing permissions |
| Escaped errors: a wrong action whose effect left the product (a customer, a filed report, the public) before anyone caught it | Incident log | Zero | up to standing permissions |
| Highest consequence tier of the action type | The tier map | Tier 1 or 2, if you chose to sign them at launch. Tier 3 only inside an envelope written in code. Tier 4 is never granted | up to standing permissions |
| Plans approved without edits | Plan approval records | At least 20 plans, 18 of them unedited | up to plan approval |
| Runs that touched something outside the plan’s declared scope | Run logs compared with the plan | Zero in those 20 | up to plan approval |
| Interrupt and resume | A test you run | One run stopped midway and resumed with no step redone or duplicated | up to plan approval |
| Evaluation score on the task class | The evaluation suite | A score with its sample size, on at least 100 cases | up to plan approval |
| Error rate on a held-out set | A labeled set the agent was never tuned on | At most 3 errors in 300 cases | up to monitored |
| Budgets on steps, time and money | The harness configuration and a test | Each budget has stopped a test run | up to monitored |
| Kill switch | A drill | Tested before the first unattended run | up to monitored |
| After-the-fact sample | Review records | 20 finished runs a week, read by a named person, four weeks running | up to monitored |
| What the task can touch | The tier map | Tier 1 and 2, plus tier 3 inside its coded envelope; any other tier 3 step and every tier 4 step stops the run and waits | up to monitored |
| Model upgrade (the book’s trigger) | The model version field in the record | Any change of version | down one |
| Escaped error or near miss on this task | Incident log | One | down one |
| New tool, wider credentials, or a new tier 3 or 4 action in the task | Change log | Any | down one |
| Edit, rejection or reversal rate against the value at promotion | The same records | Double | down one |
| Sample left unread, or approvals faster than a card can be read | Review records, timestamps | Two periods running | down one |
Why these numbers, and how soft are they?
The numbers are soft, and the sample sizes are where the softness shows. Zero problems in 200 signed actions still leaves a 95% interval that reaches 1.8%. Four edits in 200, the 2% line, gives an interval from 0.5% to 5.0%.
The plan row is weaker still. Eighteen unedited plans out of 20 is 90%, with a 95% interval from 68% to 99%. Three errors in 300 held-out cases is 1%, with an interval from 0.2% to 2.9%. The eval sample size calculator computes these intervals for your own counts.
That is why “three clean runs” proves almost nothing: zero failures in three runs is consistent with a failure rate anywhere up to 71%. The promotion decision should quote the interval, and the owner should accept its upper end in writing.
“Tier 4 is never granted” follows the chapter’s “every time” and matches the consequence tier classifier, whose dial selector shows the gate policy for each tier at each position. On tier 3 that tool and the gates post are stricter than this table: past the first position they keep an externally visible action in a review queue. The coded envelope is this post’s own further step, and a team can stop at the queue.
The “down one” on a model upgrade is my operational reading of “re-earn.” The book says the evidence expires; dropping one notch until the rows for that move are met again on the new version is one way to act on that, and a strict one. A second model that grades the first one’s output, the evaluator-optimizer pattern, can supply part of the evidence at the loose positions, provided its own miss rate is measured and recorded.
What does the autonomy slider AI agents ship with look like to a user?
Seen from the user’s chair, each position is a different screen: a card to sign, a list of permissions, a plan, or a budget with a stop button. The book supplies the frame, “the same instrument seen from the product side,” and the range, from a person in the loop to a person on it.
The glossary’s definitions are the free reference: “in the loop means inside the run, approving actions as they happen; on the loop means above it, reviewing what an autonomous run proposes and produces” (human in the loop / on the loop).
The table below is this post’s own construction. Its last column lists what the product records once a task is at each position. The evidence for the next move is gathered there too, as a trial: plans are written and read before plan approval is granted, and samples are read before a run goes unattended.
| Position | What the user sets | What the user is shown | What the user can stop or undo | What the product records |
|---|---|---|---|---|
| 1. Every consequential action signed | Which tasks the agent may attempt | One decision card per consequential action | Approve, reject, or reject with edits, per action | Signed, edited and rejected counts per action type; reversals; escaped errors |
| 2. Standing permissions | Per action type: always allow, always ask, never; caps and quotas on the allowed types | The escalations at the envelope’s edge, the always-ask types, a log of what ran | Revoke a type; undo a logged action | What ran under each permission; escalations; reversals; escaped errors |
| 3. Plan-level approval | The plan: scope, steps, budget | The plan before the run, progress during it, the result with its evidence after | Edit or reject the plan; interrupt and resume the run | The plan as approved, edits made, anything touched outside its scope, interruptions |
| 4. Monitored autonomy | Budgets on steps, time and money; the schedule; the review sample | Alerts, budget stops, a sample of finished work | The kill switch; undo of what the run changed | Budget stops, alerts, the sample read and what it found, the held-out error rate |
What do shipped products expose today?
Shipped agent products expose this instrument as a category of control usually called permission modes or approval settings. Two examples follow, each described only as its own documentation described it on October 7, 2026. They are examples of a category; I recommend neither, and both pages will change.
Example A, a coding agent’s documentation (Anthropic). The page describes each session mode by “What runs without asking”. Its table runs from “Reads only” through a mode that adds file edits, a planning mode, and a mode described as “Everything, with background safety checks”. A further mode is listed as best for “Isolated containers and VMs only” (documentation, read October 2026).
Example B, a second coding agent’s documentation (OpenAI). This page says its controls “come from two layers that work together”. One is the sandbox, which sets what the agent “can do technically”. The other is the approval policy, which sets when the agent “must ask you before it executes an action” (documentation, read October 2026).
Example B carries the transferable idea. In practice, the autonomy slider in shipped AI agents is two controls: what needs a signature, and what the environment physically allows. A wider sandbox lets more run unasked, and a narrower one makes a loose approval setting less dangerous.
A third pattern is the per-tool rule. The first vendor’s research page describes a consumer assistant where users “can configure permissions (e.g., always allow, needs approval, block) for each action” (Anthropic, 2026).
When that middle surface is missing, users ask for it. An issue opened on one open-source agent in October 2026 complains that most of a tool set prompts “on every single use” and that there is “no way to pre-approve these tools” (GitHub).
One caution about the word. In the recap of the talk that the book cites, the examples are product ladders, such as a search assistant’s “search -> research -> deep research” (Latent.Space, 2025). There the control selects how large a piece of work the agent takes on, while the four positions set how often a person signs.
What goes in the per-task autonomy record?
The record holds one task’s position, the evidence behind it, and the conditions that move it either way. The book explains why someone has to keep it: “The ledger lives with you, in permission files and eval dashboards, because the agent cannot keep it.”
The template is this post’s own. Its evidence fields match the promotion table row by row, and its “shown” and “can stop” fields match the product-side table. The owner and the next review are my additions; the book asks for neither.
AUTONOMY RECORD: [task class]
Touches (highest tier): [1 read-only | 2 reversible, logged |
3 externally visible | 4 irreversible or high-stakes]
Position now: [1 every consequential action signed |
2 standing permissions | 3 plan-level approval |
4 monitored autonomy]
Set on: [date] Model version: [...] Owner: [name]
Envelope (position 2 and looser)
always allow: [action types, with caps and quotas]
always ask: [action types; every tier 4 type is here]
never: [action types]
Evidence on file, this model version
signed actions: [n] edited or rejected: [n, %]
reversed after running:[n, %]
escaped errors: [n]
plans approved: [n] unedited: [n] outside declared scope: [n]
interrupt and resume: [tested on (date) | not tested]
evaluation score: [x of n cases]
held-out error rate: [x of n cases, with its interval]
budgets: steps [..] time [..] money [..]; stopped a test run on [date]
kill switch: [tested on (date) | not tested]
after-the-fact sample: [n per week, read by (name)]
What the user is shown: [card | escalations and log | plan, progress, result |
alerts, budget stops, sample]
What the user can stop: [reject or edit | revoke a type, undo |
interrupt and resume | kill switch]
Moves up one notch when: [the rows of the promotion table for the next move,
with your own values]
Moves down one notch when: [model upgrade | escaped error or near miss |
new tool, credentials or tier 3-4 action |
rate doubles | sample unread or approvals too fast to read]
Next review: [date, or the trigger that forces one]
How does a bookkeeping product place three tasks?
It places them at three different points, and only one of them moves today. The product is an illustrative one: an agent inside a bookkeeping tool for small firms, with 90 days of history on one model version. All numbers are invented for the example.
Categorizing transactions. A category is an internal field that can be changed back, and every change is logged, so the task is tier 2. The tier rule would let it run logged from the first day. This product had the user confirm every suggestion at launch anyway, to build a record for a new task class, and the first five rows of the promotion table are how that elective signature is retired. The records show 2,400 suggestions confirmed, 36 of them edited or rejected (1.5%), and 12 of the 2,364 accepted ones corrected later (0.5%). No wrong category reached a filed report.
Run that through the first five rows of the promotion table: 2,400 is above 200, 1.5% is under 2%, 0.5% is under 1%, no error left the product, and the tier is 2. The task moves up to standing permissions. The envelope allows known vendors under an amount cap and always asks for new vendors and split transactions. The next step is to collect the plan-level evidence for a month-end run.
Drafting and sending payment reminders. Drafting is reversible. Sending puts a message in a customer’s inbox, so the task is tier 3. The records show 140 sends signed, 21 edited and 3 rejected: 24 of 140, or 17%. All three rejections were for invoices already paid.
The sample is under 200 and the rate is far above 2%, so the task stays at position 1. The next step is a fix before any recount: a check in code that the invoice is still unpaid, and a narrower action type (first reminder, fixed template). The count restarts from zero on that narrower type.
Paying supplier invoices. A payment moves money, so it is tier 4, and the table says tier 4 is never granted. The task stays at position 1 on any evidence. The next step makes the signature cheaper to give well: one card per payment run that shows the list, the count and the total. The matching of invoices to purchase orders gets its own record as a separate tier 2 task, and that one can move.
Here is the middle task’s record, shortened to the fields that decide it.
AUTONOMY RECORD: payment reminders (draft and send)
Touches (highest tier): 3 externally visible
Position now: 1 every consequential action signed
Evidence on file: signed 140; edited or rejected 24 (17%);
escaped errors 0
Moves up one notch when: unpaid check in code; 200 signed first reminders;
at most 2% edited or rejected; at most 1% reversed;
escaped errors 0; envelope written in code
(template, amount cap, daily quota)
Moves down one notch when: already at the tightest position
Next review: after 200 signed first reminders, or on model upgrade
How do the four positions relate to published levels of autonomy?
They are related by subject and differ in what they grade, so no clean mapping exists. The published scales grade a system or an agent as designed. The book’s positions grade one task at one moment and expect it to move.
The oldest source here is Parasuraman, Sheridan and Wickens, “A model for types and levels of human interaction with automation” (IEEE Transactions on Systems, Man, and Cybernetics, Part A, 2000). It proposes four classes of function, from information acquisition to action implementation, each automated “across a continuum of levels from low to high”. Its key sentence: “A particular system can involve automation of all four types at different levels” (PubMed record).
For agents, Feng, McDonald and Zhang’s “Levels of Autonomy for AI Agents” (arXiv:2506.12469v2, 2025) defines “five levels of escalating agent autonomy, characterized by the roles a user can take when interacting with an agent: operator, collaborator, consultant, approver, and observer.” The authors treat the level as “a deliberate design decision, separate from its capability” (arXiv).
The names invite a one-to-one mapping that does not hold. At their approver level, “the user is only required to interact with the agent when the agent encounters a blocker it cannot resolve on its own.” That resembles the book’s escalation more than its signed first position. Their observer level “comes with no means for, user involvement,” while the book’s loosest position keeps a kill switch.
Mitchell, Ghosh, Luccioni and Pistilli (arXiv:2502.02649v3, 2025) grade a different axis: how much of a program’s flow the model controls, in five levels from “Simple processor” to “Fully autonomous agent”. Their claim is that “The more control a user cedes to an AI agent, the more risks to people arise” (arXiv).
None of the three papers says what a team must measure before it changes a setting, which is the gap the promotion table tries to fill.
Where does this instrument break?
It breaks where the evidence is thin, where the action cannot be undone, and where nobody owns the record. Each limit below applies to the post’s own tables more than to the book’s four positions.
The thresholds are mine. No published study that I found says how many clean actions justify a standing permission. A team with real incident costs should derive its own from the loss it can accept.
The ladder may be too strict. The book does not say a task must pass through every position. A single-step task such as categorization has little use for a plan, and a team might reasonably treat positions 2 and 3 as one stop for it. I kept one notch at a time because it forces each move to be argued separately.
Low volume never reaches the sample. A task that runs five times a month will take years to produce 200 signed actions.
Records can be gamed by habit. An approval record counts clicks, and if the signatures were reflexive, a low rejection rate measures the reviewer. Seeded requests, described in the gates post, are the check on that.
The public evidence is narrow. The usage study is one vendor’s product, used mostly for software work, and the human-factors research predates agents. Nothing here shows that a per-task record lowers incident rates.
One instrument, one record per task
The autonomy slider for AI agents is one instrument with two names and four positions, and its unit is the task. Holding an agent to it takes three things written down per task: the position, the evidence, and the events that move it down. Chapter 25 states the bet behind every setting: “the speed of verification is what the slider’s loose settings are purchased with.”
The chapter that builds the dial ends on the part no position changes: “What has not changed, at any position of the dial, is whose signature it is.”
Chapter 12, “Oversight and Autonomy,” builds the gates, the review discipline and the dial, and Chapter 25 carries the slider across eight recipes (both in the full book; the glossary entries linked above are free). The agent patterns guide places this post among its neighbors, or you can see the formats.
Questions readers ask
- What is the autonomy slider for AI agents?
- It is a control for how much an agent may do between moments of a person's attention. The book calls the engineer's version the autonomy dial and gives it four positions: every consequential action signed, standing permissions, plan-level approval and monitored autonomy. Chapter 25 uses autonomy slider for the same instrument seen from the product side, and credits the term to a 2025 talk by Andrej Karpathy.
- What is human in the loop AI, and how is on the loop different?
- The book's glossary defines both. In the loop means inside the run, approving actions as they happen. On the loop means above it, reviewing what an autonomous run proposes and produces. The two are the ends of the dial's range, and the two middle positions (standing permissions and plan-level approval) sit between them.
- How much autonomy should the first version of an agent product ship with?
- The book's rule is to start every new task class tight. In practice that means each task ships with a person signing its consequential actions, plus a written statement of what would have to be measured before the task moves one notch looser. Read-only and logged reversible actions run without a signature from day one, unless a team chooses to sign them while it builds a record.
- When can an AI agent run unattended?
- The book describes monitored autonomy as running unattended against budgets, with work reviewed after the fact, monitoring, and a kill switch within reach. This post adds measurable conditions: budgets that have stopped a test run, a tested kill switch, a held-out error rate with its sample size, a named person reading a sample, and a task whose irreversible steps, and externally visible steps outside a coded envelope, stop the run and wait.
- Does a model upgrade change the autonomy setting?
- Yes. The book says the evidence expires with the model version it measured, and tells you to re-earn the dial settings after every upgrade. This post's working rule is to drop each task one notch on a model change and move it back when the promotion table's rows for that move are met again on the new version.
Sources
- Latent.Space (2025). Andrej Karpathy on Software 3.0 (recap of the talk “Software Is Changing (Again)”, page dated 17 June 2025)
- Geoffrey Litt (2025). Code like a surgeon
- Raja Parasuraman, Thomas B. Sheridan, Christopher D. Wickens (2000). A model for types and levels of human interaction with automation, IEEE Transactions on Systems, Man, and Cybernetics, Part A, 30(3):286–297, doi:10.1109/3468.844354
- Raja Parasuraman, Dietrich H. Manzey (2010). Complacency and bias in human use of automation: an attentional integration, Human Factors 52(3):381–410
- K. J. Kevin Feng, David W. McDonald, Amy X. Zhang (2025). Levels of Autonomy for AI Agents, arXiv:2506.12469v2 (28 July 2025); Knight First Amendment Institute essay series
- Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, Giada Pistilli (2025). Fully Autonomous AI Agents Should Not be Developed, arXiv:2502.02649v3 (20 October 2025)
- Anthropic (2026). Measuring AI agent autonomy in practice (18 February 2026)
- Anthropic (2026). How we contain Claude across products (25 May 2026)
- Anthropic (2026). Trustworthy agents in practice
- Anthropic (2026). Choose a permission mode (product documentation; read 7 October 2026)
- OpenAI (2026). Agent approvals & security (product documentation; read 7 October 2026)
- lielcohen (2026). Hacker News comment 47446330 (19 March 2026), with a reply by codingdave
- kissgyorgy (2026). Hacker News comment 48767887 (2 July 2026)
- bruckie (2026). Hacker News comment 48404757 (4 June 2026)
- thefounder (2026). Hacker News comment 49896956 (29 September 2026)
- rglover (2026). Hacker News comment 48053853 (7 May 2026), with a reply by kelseyfrog
- natesilva (2026). osaurus issue 3025: approval prompt on every use of a tool (opened 6 October 2026)