To design agent tools and schemas for software built for humans, choose the surface one operation at a time. Take the highest rung that exists, completes with nobody present and can be read back. Wrap it behind a five-part contract: inputs, output, errors, side-effect class, timeout. The model sees one task-shaped tool, and the old interface stays inside the wrapper.
I wrote this for a backend engineer whose agent must operate something nobody will rewrite: a vendor’s admin panel, a desktop accounting program, an old command-line program. Someone in the meeting said “just use computer use”. Here is an order to try surfaces in, a contract and a checklist.
Why was your software built for a person, and what does that cost an agent?
Software built for a person keeps part of its interface in that person’s head, and an agent has no access to that part. The person answers the prompt, remembers the help text and spots the duplicate row. Take the person away and each of those becomes a failure to design for.
The book argues this in Chapter 5, “Tools and the Action Space” (in the full book): “But much of what an agent drives was built years before agents existed: command-line programs, internal services, the ordinary machinery of your systems.” Those programs assume what the chapter calls “a forgiving, resourceful user”. The bill for removing her arrives in one sentence: “Every gap that human resourcefulness papered over becomes a hang, a burned retry, or a duplicate record.”
Chapter 5’s remedy is to “make explicit what a human used to supply implicitly”, and a wrapper is where you do it. In the book’s glossary a tool is “A function the model can ask your program to run, described to it as a name, a description, and an argument schema; the function itself never leaves your process.” The wrapper is that function, and what the old program expected from a person now lives inside it.
What is the book’s three-rung ladder?
The book ranks three ways to reach another system: an API first, a structured browser tool second, pixel-level computer use last. Its advice follows: “All of which arranges itself into a ladder, and my advice is to descend it only as far as you are forced.” The ranking and the reasons below are the book’s.
- API. “An API is a deterministic contract; use it wherever one exists.”
- Structured browser. A tool “driving the page through the DOM and accessibility tree” comes next, “still routed through a model’s judgment, but targeting by meaning instead of by position”. The accessibility tree is the page’s simplified form for screen readers: each interactive element with a role, a name and a state.
- Pixels. Here the model gets a view of the screen, a pointer and a keyboard. The chapter calls this rung “maximally general, slowest, most fragile, reserved for software that offers no better door”.
On how to live on the ladder, the chapter says: “The strongest pattern treats the lower rungs as reserve rather than residence; the agent works through APIs for, say, nine steps in ten, and drops to the screen only for the gap.” The nine in ten is an illustration, flagged by the book’s own “say”.
How do you design agent tools and schemas across six surfaces?
To design agent tools and schemas for an existing system, walk one ladder of six surfaces from the top, once per operation. The book ranks rows 1, 5 and 6; it calls the command line a “programmatic door”, gives it retrofit rules and never ranks it against an API, so row 2 is my placement. File exchange and recorded UI scripts are absent from the chapter, so rows 3 and 4 are mine. I also stretch row 5 from the browser to any program whose controls expose a role and a name to assistive technology.
| Surface (whose rung) | What the model gets | How it fails | Step down when | Steps chosen by |
|---|---|---|---|---|
| 1. Documented API (the book’s top rung) | A task-shaped tool over one or more API calls | Coverage gaps; oversized responses | The operation is missing from the API, or no credential is available for it | Wrapper code |
| 2. Command-line program (this post’s placement) | The same tool; the command line stays in the wrapper | Hangs on a prompt, a pager or a credential question; output drawn for eyes | The program cannot finish with nobody present, or its output cannot be parsed | Wrapper code |
| 3. File exchange (this post’s addition) | A receipt or the matching rows; the file stays out of the context | Stale data; partial imports; a renamed column; a person has the file open | The system reads or writes no file for this operation unless someone is at its screen | Wrapper code |
| 4. Recorded UI script (this post’s addition) | The same tool; a fixed sequence over named elements runs inside | A redesign or a new dialog breaks it; a click changes nothing | The elements carry no role or name, the path differs from run to run, or the script breaks faster than you repair it | Wrapper code |
| 5. Structured UI driven by the model (the book’s middle rung) | Actions on elements by role and name, plus a bounded observation | Snapshots flood the context; clicks without effect | The elements carry no role or name: a canvas, a remote desktop, custom-drawn controls | The model |
| 6. Screen (the book’s bottom rung) | A small generic action set: look, point, type, scroll | Slowest; an image per step; missed coordinates | No lower rung exists: keep a person on it, or ask for a better door | The model |
Filter on the last column: on rows 1 to 4 your code chooses every step and the model sees one task-shaped tool; on rows 5 and 6 the model chooses the steps and sees element or screen actions, under the same contract.
What does each surface look like in practice?
The two specimens per row are dated, and none is a recommendation.
- Documented API. One provider’s “nearly 3,000 HTTP API operations” behind a single schema (Cloudflare blog, 13 April 2026); the generated endpoints in the Reflex comparison below.
- Command-line program. A code host’s official client, the baseline of the Scalekit benchmark below; a client for an office-suite API with published rules for agent callers (Poehnelt, 4 March 2026).
- File exchange. The File Transfer pattern of Hohpe and Woolf (read 7 October 2026); a paper on why a raw spreadsheet is too large to hand a model (Dong et al., arXiv, July 2024).
- Recorded UI script. I found no primary documentation, only accounts of robotic process automation, which replays a person’s clicks. One Hacker News commenter (30 July 2026): “Usual RPA is too brittle and maintaining it often requires more work than just doing the work”.
- Structured UI. A browser-automation server working “through structured accessibility snapshots” (Microsoft’s playwright-mcp README, read 7 October 2026); a production browser agent (Vardanyan, arXiv, November 2025).
- Screen. A capability its own lab says “remains slow and often error-prone” (Anthropic, 22 October 2024); a screen-control tool’s documentation (read 7 October 2026).
What do the published measurements show, and what do they leave out?
Two published comparisons put an API above a browsing or screen agent, each on its own tasks, and a third compares packagings. I found none that measures a command-line or file surface against an API.
One admin task, a screen agent against endpoints
A vendor’s write-up ran one multi-step admin task both ways with the same model (Palash Awasthi, Reflex blog, 27 April 2026). Reflex sells a framework that generates such endpoints. The screenshot-driven agent took 53 ± 13 steps and 1,003 ± 254 seconds across three runs. Tool calls to HTTP endpoints took 8 calls and 19.7 ± 2.8 seconds across five runs.
Uncached input tokens came to 550,976 ± 178,849 against 12,151 ± 27, the ratio behind the title’s “45x”. The screen agent failed when left to find its own way. These are its numbers after the authors wrote “Fourteen numbered instructions” for it, on one library version and a small pinned dataset.
API calls against browsing on a web benchmark
A research paper compared a browsing agent, an API-calling agent and a hybrid of both (Song, Xu, Zhou and Neubig, arXiv:2410.16464, submitted 21 October 2024). The benchmark was WebArena, a set of web navigation tasks, and the abstract states that “API-Based Agents outperform web Browsing Agents” and that the hybrid gained “more than 24.0% absolute improvement over web browsing alone, achieving a success rate of 38.9%”.
That 38.9% is the share of WebArena tasks the hybrid completed, in one paper, with the models of its date. The order is the part I would carry forward, as one paper’s result on one benchmark: the agent with both doors did best.
Two packagings of one API
A second vendor benchmarked a shell with a command-line client, the same with a tips file, and a remote protocol server (Ravi Madabhushi, Scalekit blog, 11 March 2026). Scalekit sells agent authentication and ran it itself: 75 runs, 25 per arm, five read-only tasks on one repository. Completion was 25 of 25, 25 of 25 and 18 of 25, and of the seven misses it says: “Every failure was a TCP-level timeout”. Median tokens per run were 4 to 32 times higher through the protocol server than through the bare command-line arm, depending on the task.
Every arm reaches the same API, so the benchmark is silent on surfaces.
What did nobody measure?
No source I found measures a command-line program against an API, or a file exchange against anything. Rows 2, 3 and 4 are argued from properties: who parses the output, what can hang, whether a model turn is spent per step.
How do you choose the starting rung for one operation?
Walk the ladder from the top for one operation and start at the first rung that passes three tests. If no rung passes, a person keeps the operation.
- Present. The surface exists for this operation; the row’s “Step down when” cell says what absence looks like.
- Unattended. The whole operation completes with nobody present, and wrapper code can answer every question the surface asks. A second factor mid-run fails this test.
- Readable back. The wrapper can read the resulting state, through this rung or a higher one. A read leaves no new state, so for a read the test is that the output parses into the named fields, and
verified_byrecords that parse.
The walk stops at the first pass, so an operation gets exactly one starting rung. An operation is the smallest action with one effect; the read that verifies it is a separate operation with its own walk.
Why per operation?
One system can sit on several rungs at once, so you design agent tools and schemas per operation. A ticket system may serve every read on rung 1 and lack the call that closes a ticket with a note. That operation walks down to rung 4 while the reads stay put: “reserve rather than residence” in practice.
One shortcut gets no rung here: calling the HTTP endpoints the page itself calls, a mechanism one Hacker News commenter (6 May 2026) says “can usually be reverse-engineered and easily emulated”. That is an undocumented API: it can change without notice, and using it may breach the system’s terms.
When do you step down, and when do you climb back?
Step down on evidence: the failure in the row’s “Step down when” cell, seen in your transcripts more than once. The operation moves to the next rung below that passes the three tests.
Climb by repeating the walk from the top when the wrapper breaks, when the system’s owner ships a release, and on a fixed date. A higher rung that now passes takes the operation.
How do you wrap a command-line program so it cannot hang?
Run the program with no terminal and an input the wrapper owns. Fix every argument in code, put one deadline on the call, and give every ending a named result.
The chapter’s first retrofit rule addresses the program’s author: “Never block on an interactive prompt: every input should be passable as a flag, an argument, or a file”. Its reason is that “the silent hang is the most expensive failure an agent can hit, because it learns nothing and recovers slowly”. When you cannot change the program, the wrapper carries the rule.
What do you do about prompts, pagers and credential helpers?
Decide in code what each question gets. An issue on one coding agent’s tracker (1 June 2026) states the mechanism: a child process that “blocks on stdin will hang indefinitely”. It also proposes a remedy: “Closing stdin causes interactive prompts to fail fast (EOF) instead of hanging the session”. Four wrapper rules follow.
- Inputs. Supply every input as an argument, a file or an environment value.
- Known confirmations. The wrapper writes the answer once, after your gate approved the call. The confirmation a person used to give moves to the contract’s gate.
- Anything else. The wrapper kills the program and returns
PROMPT_UNEXPECTEDwith the question as the hint. A credential question returnsAUTH_REQUIREDand never reaches the model. - Pagers and helpers. Turn them off through the program’s settings or environment. One report (20 February 2026) has commands that “still attempt to use an interactive pager”; another (26 June 2025) has a signing step asking for a password.
The deadline stays in every case. A report on a second agent’s tracker (5 April 2026) describes a command waiting for credentials: “It just hangs silently until the 5-minute timeout is reached”.
Should the model see the command line or the help text?
No. A pass-through tool with one free-text argument hands the model the whole human interface. The chapter prices the habit: “It succeeds slowly, after extra help-reads and retries, and that tax is charged on every invocation for the life of the tool.” Once the wrapper fixes the arguments, thin help stops mattering.
I ran both shapes through the site’s tool schema linter on 7 October 2026. A pass-through definition, run_inventory with one string argument, returned 9 findings: 1 to fix first, 5 that should be fixed, 3 to consider. A task-shaped definition with typed, bounded arguments, a request id and a dry-run switch returned 0 findings.
Shaping names, descriptions and arguments is its own subject: how to design tools for LLM agents when you control both ends.
How do you handle exit codes and output size?
Map every exit path to a stable error code, and cap what comes back. The chapter wants programs to “Emit machine-parseable output (JSON is the common choice) with results on standard output and diagnostics on the error stream”. The wrapper parses the first stream into fields and keeps the tail of the second for the hint.
A zero exit status proves only that the program ended, so the wrapper reads the state back before it reports changed. On size the chapter’s rule is “Return the relevant slice, never the full dump”, with a notice when you cut, “because a silent cut reads as a complete answer”.
run_wrapped(program, fixed_args, inputs, deadline):
args = fixed_args + validate(inputs)
child = start(program, args, terminal: none, input: a pipe the wrapper owns,
env: paging off, color off, credentials from the wrapper's store)
until child exits or deadline passes:
keep at most CAP characters of each stream
if output stalled on a known confirmation and the gate approved: write the answer, once
else if output stalled on any other question:
kill child; return PROMPT_UNEXPECTED or AUTH_REQUIRED, hint: the question
if deadline passed: kill child and its children
return TIMEOUT, outcome: not_started if nothing was sent, else unknown
if exit status is not zero: return the error from the exit-status table, hint: error-stream tail
fields = parse(standard output) or return SURFACE_CHANGED
return fields + effect from the read-back
What about a create the program cannot deduplicate?
Keep the missing key in the wrapper. The chapter’s rule is that “a retried create should return the existing record, flagged as pre-existing, instead of minting a second one”. An old program remembers no request, so idempotency is retrofitted around it.
The wrapper keeps a ledger from a request_id to the record it created, and a repeat returns effect: existing. After a timeout the wrapper cannot know whether the write landed, so it returns outcome: unknown with the read that settles it.
How do you wrap a user interface that floods the context and reports false success?
Return a bounded observation, and verify every action by reading state back.
How do you keep the observation small?
Decide in the wrapper which fields the operation needs and return only those. One issue on a browser-automation server (10 May 2025) opens: “the snapshot returned is too long for some website”.
On rung 4 the model sees none of the page. On rung 5 the wrapper returns the named fields the operation cares about and a capped list of actionable elements. The full snapshot goes to a file, and the tool returns its path with truncated: true.
How do you know the action took effect?
Read one named piece of state before the action and again after it, and report the difference. An issue on a browser-agent library (1 October 2026) documents the false success. A click on a disabled button “returns a success result”, although “The browser never dispatches clicks to a disabled control, so nothing happened”.
An earlier issue in that tracker (5 December 2025) reports of a download button: “It Downloads the file multiple times”. The chapter covers both: “The harness should wait for and confirm state changes before the next action, and the model should look again after anything consequential; clicking and hoping is how these agents fail silently.”
act_and_verify(action, expected, read_state, wait):
before = read_state() # one named field
perform(action)
poll read_state() until it matches expected or differs from before, at most wait seconds
after = read_state()
if after matches expected: return effect: changed, observation: summary(after)
if after equals before: return effect: no_change, hint: what appears to block the action
return error UNEXPECTED_STATE, observation: summary(after)
read_state should use the highest rung available: after a scripted click closes a ticket, read the ticket through the API.
Can a file or an export be the interface?
Yes, when the system reads an import or writes an export on its own, with nobody at its screen: a watched folder, a scheduled export, a data file in a documented format. The wrapper handles the file, and the tool returns a receipt or the matching rows.
Data can be stale, so the tool returns the export’s timestamp as as_of. A column can be renamed, so the wrapper checks the header first and returns SURFACE_CHANGED when it differs. An import can be partial, so the wrapper counts records before and after. A person can have the file open, so the wrapper writes only while it holds the lock.
Cap the rows returned and leave the rest on disk for a program to filter; in the book’s example of a ten-thousand-row export, “the rows stay in the runtime”. An export behind a menu is a UI operation whose result happens to be a file, and it starts on rung 4 or lower.
What goes in the wrapper contract?
Five parts: inputs, output, errors, side effect, timeout and budget. A heading records the rung and a closing line the session. When you design agent tools and schemas for a system you cannot change, fill this in before any wrapper code.
WRAPPER CONTRACT
Tool <name the model sees> · <the operation in one sentence, one effect>
Rung <1 API | 2 command line | 3 file | 4 recorded script | 5 structured UI | 6 screen>
ruled out <each higher rung and the test it fails: present, unattended, readable back>
step down <the observation that moves this operation to the next rung below>
climb <the observation that moves it back up>
1 INPUTS <name> · <type> · <allowed values or pattern>; validated before the system is touched
never an input: a credential; on rungs 1 to 4, also a command line, a selector, a coordinate
2 OUTPUT fields: <what the next step needs, plus handles such as id, path, url>
effect: changed | existing | no_change (reads: none)
verified_by: <what the wrapper read back, and through which rung>
cap: <maximum rows or characters>; when cut: truncated: true + how to narrow
3 ERRORS <CODE> · retry: yes | no | after_check · hint: <a correct call, what to try next>
always present: AUTH_REQUIRED, SURFACE_CHANGED, TIMEOUT
4 SIDE EFFECT class: read | reversible | visible to others | irreversible
gate: none | log | review | signature (decided outside the model)
repeat: <what a second identical call does, and the key that makes it identical>
5 TIMEOUT deadline: <seconds> · budget: <maximum steps on rungs 4 to 6, maximum attempts>
on expiry: stop; return TIMEOUT with outcome: not_started | unknown
settle: <the read that tells which, before any retry>
Session <who signed in, where the credential lives, what the tool returns on expiry>
Inputs follow the chapter’s rule to “Validate as strictly as you would for a public API, because the caller, however capable, is not a trusted operator”. Errors follow its formula: “what happened, what a correct call looks like, what to try next”.
The four side-effect classes and their gates are the tiers of the site’s consequence tier classifier. I would not let rungs 5 or 6 perform an irreversible action without a person approving that step.
What should you check when you wrap one operation?
Work through the list for one operation, in order. An item that names a rung applies only when the operation starts there.
- Name the operation as one task with one effect.
- Walk the ladder from the top; write the starting rung and the test each higher rung fails.
- Write the step-down observation from the table’s row, and the one that lets it climb.
- Fix every argument that does not vary in wrapper code; validate the rest first.
- Run the call with nobody present and confirm it returns.
- Rung 2: answer or remove every prompt; turn off paging. Rungs 4 to 6: wait for named states and cap the steps.
- Parse the result into named fields, cap its size, add the truncation notice.
- Add the read-back; keep
no_changeandexistingdistinct fromchanged. - Map every ending to an error code with a retry rule and a hint.
- Classify the side effect; attach the gate outside the model; make a repeat safe.
- Set the deadline; write the
not_startedandunknownpaths and the read that settles them. - Keep the credential out of the model’s context; return
AUTH_REQUIREDon expiry. - Call the create twice with one request id; confirm the second returns
existing. - Lint the definition, run three real tasks, and read the transcripts whole.
Worked examples: where do three real operations land?
Three operations walked through the rule land on two rungs, each with one condition that sends it down and one that brings it back. The systems are fictional and the numbers illustrative.
| Operation | Starting rung | Ruled out above | Steps down when | Climbs back when |
|---|---|---|---|---|
| Close a ticket with a resolution note; the API lacks that call | 4, recorded script; read-back on rung 1 | 1: call absent. 2: no command-line client. 3: no import changes status | A ticket type shows a field the script does not know: rung 5 | The vendor adds the call: rung 1 |
| Export last month’s invoices from a desktop accounting program; menu and export dialog only | 4, recorded script; read-back by opening the file | 1, 2: neither exists. 3: the export sits behind the menu | A month shows an extra dialog: rung 5 | A scheduled export appears: rung 3 |
| Restart a service and confirm it is healthy; the internal program prompts and pages | 2, command line; read-back on rung 2 | 1: no API | The program demands a terminal the wrapper cannot script, or its status stops parsing: rung 4, the dashboard’s button | The owners expose an API: rung 1 |
The accounting row rests on a fact to check first: an accessibility inspector lists the program’s menu and dialog fields by role and name. If it lists nothing, the walk continues to rung 6. A terminal-screen program follows the same check: if the wrapper can script the session and read fields at fixed positions it is rung 2; if not, rung 6.
Closing a ticket with a resolution note
Reads stay on the API and only the close walks down.
ticket_close_with_note rung 4 recorded script · read-back on rung 1
ruled out 1, 2, 3 fail "present"
step down repeated SURFACE_CHANGED on one ticket type -> rung 5 for that type
climb the vendor's release notes list the call -> rung 1
1 inputs ticket_id (pattern T-000000) · resolution_note (1 to 2,000 characters) · request_id
2 output ticket_id, status, closed_at, url · effect · verified_by: API read · cap 600 characters
3 errors NOT_FOUND · CLOSED_WITH_OTHER_NOTE · AUTH_REQUIRED · SURFACE_CHANGED · TIMEOUT (after_check)
4 side effect visible to others (the customer is notified) · gate: review
repeat: the same ticket_id and note returns effect: existing
5 timeout 45 s · 8 script steps, 1 attempt · settle: ticket_get on rung 1
on expiry: not_started if the page never loaded, else unknown
session a service account, signed in by a person in the wrapper's browser profile
Restarting a service and confirming it is healthy
The wrapper answers the one known confirmation after the review gate approves, turns paging off, and polls the status until the service is healthy. A step down would put an internal system on a UI rung, which is the first stop case below.
service_restart rung 2 command line · read-back on rung 2 (a status read)
ruled out 1 fails "present"
step down repeated PROMPT_UNEXPECTED or SURFACE_CHANGED -> rung 4, the dashboard's button
climb the owners expose restart and health as an API -> rung 1
1 inputs service (enum of allow-listed names) · environment (enum) · reason · request_id
2 output service, state_after, healthy, started_at · effect · cap 800 characters
verified_by: status read; started_at later than the call
3 errors BAD_INPUT · PROMPT_UNEXPECTED · UNHEALTHY_AFTER_RESTART (no retry)
AUTH_REQUIRED · SURFACE_CHANGED · TIMEOUT (after_check)
4 side effect visible to others (users see an interruption) · gate: review
repeat: a ledger keyed by request_id returns the first result, effect: existing
5 timeout 120 s · 1 attempt · settle: service_status, compare started_at with the call
on expiry: kill the child; not_started if the confirmation went unanswered, else unknown
session the wrapper's own service credential, from its secret store
Who holds the login and the session?
The wrapper holds the session, and the model never sees the credential. A person signs in once, inside the browser profile or desktop session the wrapper uses, and a second factor is a handoff to that person. When rung 1 was ruled out for lack of a credential, a lower rung under a person’s session acts as that person, so set the gate by the side-effect class before the first run.
An issue on a browser-automation server (20 June 2025) opens with a window that appears “but it is not logged in”. The maintainers reply that the server keeps its own profile, and the person signs in there. On expiry the tool returns AUTH_REQUIRED.
The glossary’s entry for sandbox gives the reason to keep credentials away from the model: “a credential never mounted inside the boundary cannot be read by any injection, because the target is absent”. Any page the agent reads is input written by strangers, which is how prompt injection arrives. Isolation and credential brokering belong to sandboxing an agent’s tool execution.
CLI, protocol server or plain function: which package?
Packaging is a separate decision from the surface, and the contract is the same under each. A wrapper can reach the model as a plain function, a shell command or a protocol server; the Scalekit benchmark measured that choice. The decision between a tool protocol and plain function calling turns on how many agents and systems share the tool, and on who runs the other end.
When should you stop retrofitting and ask for an API?
Stop in three cases: the operation has reached a UI rung on a system your own organization owns; it has stepped down twice; or its effect is irreversible and its only door is a UI. In each case the wrapper costs more than a request to the system’s owner.
A Hacker News commenter (6 May 2026) wrote of screen control: “It is your last resort. Do not use it on state that lives in a DB that you own.” Ask the owners for an endpoint, or for the chapter’s retrofit rules applied to the program itself.
In source code, making a legacy codebase legible before a coding agent works in it is the same step pointed at files: write down what a person used to know.
What are the limits of this guide?
This guide measures nothing itself. Its numbers are other people’s, and two of the three comparisons were written by vendors.
- Argued placements. Rows 2, 3 and 4, and the desktop stretch of row 5, rest on reasoning. No controlled comparison supports them.
- Security is only named. A page is untrusted input, and the chapter’s interim posture is “least privilege, a human confirmation in front of consequential actions”.
What is the one rule to keep?
Choose the rung per operation, and make the wrapper tell the truth about what happened. To design agent tools and schemas for software you cannot change is to write down what a person used to supply: the answer to the prompt, the glance that confirmed the click, the memory of having done it already. The contract makes each of those testable.
The ladder and the retrofit rules are in Chapter 5, “Tools and the Action Space”, in the full book. The glossary entries for tool and action space are free, and the guide to tools, skills and protocols collects the related tools. When you want the whole chapter, see the formats.
Questions readers ask
- Can an AI agent use software that has no API?
- Yes, through a lower surface, chosen per operation. Check in order for a command-line program, a file the system reads or writes on its own, a user interface whose elements carry roles and names, and last the screen itself. Take the first one that exists for the operation, completes with nobody present and can be read back, then wrap it so the model sees one task-shaped tool.
- Is computer use reliable enough for production work?
- It depends on the operation, and the published evidence is dated and narrow. One vendor-authored write-up of a single admin task (Reflex, 27 April 2026) needed a 14-step written walkthrough before its screenshot-driven agent succeeded, at 53 ± 13 steps over three runs. Chapter 5 of the book calls pixel-level control the last rung, reserved for software that offers no better door. Use it for the gap, verify every effect by reading state back, and keep irreversible actions behind a person.
- Why does my agent hang when it runs a command?
- The program is waiting for a person: a confirmation prompt, a pager, or a credential question. The agent cannot answer what nobody surfaced to it. Run the program with no terminal attached and an input the wrapper owns, supply every input as an argument or a file, answer known confirmations in code after your gate approves, and put a deadline on the whole call that returns a named timeout.
- Should a command-line tool return JSON to an agent?
- Separate the two readers. The wrapper needs structure it can parse reliably and compare against a post-condition, so ask the program for a machine-parseable format when it has one. What the wrapper then returns to the model can be a few named fields in compact text, capped in size, with a notice when it was cut.
- How does a browser agent stay logged in?
- The wrapper holds the session. A person signs in once inside the browser profile the wrapper uses, the wrapper reuses that session, and a second factor is a handoff to a person. The model never receives a password, cookie or code. When the session expires, the tool returns an authentication error that says a person must sign in again.
Sources
- Palash Awasthi (Reflex) (2026). Computer use is 45x More Expensive Than Structured APIs (vendor blog, 27 April 2026)
- Ravi Madabhushi (Scalekit) (2026). MCP is up to 32× more expensive than CLI. Here's why we still use it. (vendor blog, 11 March 2026)
- Yueqi Song, Frank Xu, Shuyan Zhou, Graham Neubig (2024). Beyond Browsing: API-Based Web Agents (arXiv:2410.16464, submitted 21 October 2024; abstract read)
- Aram Vardanyan (2025). Building Browser Agents: Architecture, Security, and Practical Solutions (arXiv:2511.19477, submitted 22 November 2025; abstract read)
- Haoyu Dong et al. (2024). SpreadsheetLLM: Encoding Spreadsheets for Large Language Models (arXiv:2407.09025, submitted 12 July 2024; abstract read)
- Anthropic (2024). Developing a computer use model (22 October 2024)
- Anthropic documentation (2026). Computer use tool, section Limitations (vendor documentation, one example of a screen-control tool; read 7 October 2026)
- Microsoft (2026). playwright-mcp README (one example of a browser-automation server; read 7 October 2026)
- Gregor Hohpe and Bobby Woolf (2026). File Transfer (pattern page from Enterprise Integration Patterns; read 7 October 2026)
- Justin Poehnelt (2026). You Need to Rewrite Your CLI for AI Agents (4 March 2026)
- Cloudflare blog (2026). Building a CLI for all of Cloudflare (13 April 2026)
- huisjes (anthropics/claude-code issue tracker) (2026). Issue #64594: Bash tool inherits harness stdin by default (opened 1 June 2026)
- jyongchul (google-gemini/gemini-cli issue tracker) (2026). Issue #24707: run_shell_command hangs for 5 mins on interactive/slow commands (opened 5 April 2026)
- victoropp (anthropics/claude-code issue tracker) (2026). Issue #27136: Git commands freeze during commit/push workflow due to pager interaction (opened 20 February 2026)
- codebam (google-gemini/gemini-cli issue tracker) (2025). Issue #1689: Run blocking/long running shell commands in background (opened 26 June 2025)
- Weixuanf (microsoft/playwright-mcp issue tracker) (2025). Issue #395: the snapshot length is too long, add pagination to it (opened 10 May 2025)
- yotambraun (browser-use/browser-use issue tracker) (2026). Issue #5964: click reports success on a disabled button that has a click listener (opened 1 October 2026)
- jeel-patel-scaletech (browser-use/browser-use issue tracker) (2025). Issue #3722: Agent fails to keep track of downloads when no visual cues are present (opened 5 December 2025)
- alza-bitz, with a maintainer's reply (microsoft/playwright-mcp issue tracker) (2025). Issue #584: Logged in to sites, but the browser window opened by Playwright is not logged in (opened 20 June 2025)
- theptip (Hacker News) (2026). Hacker News comment on computer use for internal apps (6 May 2026)
- jasomill (Hacker News) (2026). Hacker News comment on emulating a site's own HTTP calls (6 May 2026)
- euphetar (Hacker News) (2026). Hacker News comment on the upkeep of recorded automation (30 July 2026)