To sandbox AI agent tool execution, run model-written code in a sealed room. The room has its own filesystem, a network denied by default, no standing credentials, hard caps on time and resources, and a fresh start for every session. Only a small result leaves, and each of those walls can be tested from inside.
I wrote this for the backend engineer who has just been asked where the agent’s code will run. By the end you can write the policy for a code tool, run sixteen tests that show each wall holds on your own sandbox, and say what changes at one sandbox per session.
The cost of skipping the room is on the public record. In March 2026 an engineer published his account of a coding agent that, running an infrastructure tool on his computer, issued a destroy command against a course platform’s production infrastructure. The database came back about a day later through the cloud provider’s support, with “1,943,200 rows in the courses_answer table” (Grigorev, 2026).
Why is code execution the one tool that stands in for the others?
Code execution stands in for the other tools because almost anything a computer can do can be written as a program. Chapter 5 of AI Agents, Engineered introduces it as the third option when a task steps off the menu of fixed tools: “give the agent one tool that writes and runs programs”.
The contract is one clause of that chapter: “the agent emits a script, an isolated runtime executes it, and the output returns as the tool result”. The chapter’s summary: “One capability composes all the rest, which is why practitioners have taken to calling code a universal interface.”
One infrastructure team wrote in 2025 that “LLMs are better at writing code to call MCP, than at calling MCP directly”, where MCP is a protocol for connecting agents to tools (Varda and Pai, 2025). The book states the division of labor in one line: “the model decides what to compute; code computes it”.
What does that one tool cost?
It costs a piece of infrastructure and an author you cannot vouch for. Chapter 5 puts the first cost plainly: “Code execution is infrastructure: a sandbox must be provisioned, secured, monitored, and paid for, and it brings cold starts and moving parts that a plain tool call does not have.”
The second cost is the reason for the room. The chapter says: “The sandbox is mandatory because of who wrote the code and what may have influenced them.” The author is probabilistic and may have read hostile text on the way, which is the subject of how to prevent prompt injection in AI agents. Hence the chapter’s rule: “The rule does not bend: never run model-written code with access to anything you would not hand a stranger.”
When the jobs are few and fixed, named tools are simpler to permission and audit, and agent tools for software built for humans are usually the first thing to build. If you design agent tools and schemas for three well-understood operations, you may never need an interpreter. The guide on how to design tools for LLM agents draws that line.
What does it take to sandbox AI agent tool execution?
It takes seven properties the book asks for and two that this post adds, and each one stops a named failure and leaves another standing. Chapter 5 defines the sandbox in one sentence: “The sandbox is the sealed room the interpreter runs inside: its own filesystem, capped processor, memory, and time, and no network except what you deliberately allow.”
One forum commenter, writing about sandbox interfaces, wanted “an honest matrix: which threats this layer stops, which ones it only slows, and which ones still need a microVM or a disposable host” (Hacker News, September 2026). The table below is my attempt. The arrangement is mine; the book states these requirements across three chapters.
| # | Property | What it stops | What it does not stop | Where the book asks for it |
|---|---|---|---|---|
| 1 | A filesystem of its own, holding only the workspace and named inputs | A bad or hijacked script reading your credentials or deleting your files | Damage to whatever is mounted writable | Chapters 5 and 17 |
| 2 | Network denied by default, destinations allowed deliberately | Code sending what it touched to an address nobody approved | Data leaving through an allowed destination, or through a link in the agent’s reply | Chapters 5, 17 and 20 |
| 3 | No standing credentials inside; any grant injected per task and scoped to it | Theft of a key that was never in the room | Misuse of the scoped grant, within its scope, while it lives | Chapters 17 and 20 |
| 4 | Hard caps on processor, memory, wall-clock time and output size | A runaway script starving the host or flooding the context window | A wrong answer delivered inside the caps | Chapters 5 and 20; caps on process count and disk are this post’s additions |
| 5 | A fresh sandbox per session and per tenant, torn down afterward | Leftovers passing from one run, or one customer, to the next | Loss of work nobody saved outside | Chapters 5 and 20 |
| 6 | A boundary enforced by the operating system or hypervisor and built from battle-tested primitives | Being talked out of a rule | A gap in the layer you wrote yourself, or a flaw in the primitive | Chapter 17 |
| 7 | A narrow return channel: exit code, capped output, any cut announced | Bulk data and intermediate values passing through the model | A result nobody validated | Chapter 5 |
| 8 | Fail closed: when the boundary cannot be set up, the tool refuses to run | A silent fall back to running on the host | A boundary that starts and is misconfigured | This post’s addition, from a 2026 fleet write-up |
| 9 | Files that leave are untrusted input to the host | A planted file being run by a person or a job outside | A reviewer who approves without reading | This post’s addition, from practitioners’ reports |
Rows 1 to 7 paraphrase the book closely. Chapter 5 puts the network rule first: “Deny the network by default and allow specific destinations deliberately, because a sandbox that can phone out can exfiltrate whatever the code touched.”
Which isolation mechanism stands between the code and the host?
Five categories of mechanism do, and each sets the thickness of one wall: the boundary between the code and the host’s kernel. Network, credentials, caps and lifecycle are separate walls that every category leaves to you.
Chapter 5 names three strengths, as examples of a category: “ordinary containers at the light end, application kernels that intercept system calls, per-session micro virtual machines at the strong end”. The first and last rows below are my additions, and each row lists two documents I read. For containers and user-space kernels I found one implementation, so the second document describes the class or a service that runs the same one.
| Category | What stands between the code and the host | Stated limit | Examples and documents read (dated) | Where the sources describe it in use |
|---|---|---|---|---|
| Process-level restriction | The host’s own kernel, applying rules on which paths, and sometimes which ports, the process may use | Same kernel as everything else on the machine; the protection is exactly the rules you wrote | Landlock, a Linux security module (kernel documentation, page dated August 2026); bubblewrap, a Linux sandboxing tool (README, read October 2026) | own machine |
| Container | A separate view of files, processes and network on a kernel shared with the host | One runtime’s security documentation says its defaults may give incomplete isolation and that only trusted users should control its daemon | Docker Engine (security documentation, read October 2026); the class as described in NIST SP 800-190 (September 2017) | own machine, single-tenant service |
| User-space kernel | An application kernel that receives the code’s system calls in place of the host kernel | Reduced application compatibility and a higher cost per system call; no protection against hardware side channels (its own documentation) | gVisor (documentation, read October 2026); the same runtime, as the default that Modal Sandboxes documents (read October 2026) | single-tenant service, multi-tenant fleet |
| MicroVM | A small virtual machine per workload, with its own kernel, behind hardware virtualization | More resource overhead than a container (one vendor’s documentation for local agent sandboxes, read October 2026); a mounted workspace is still inside the wall | Firecracker (design paper, NSDI 2020); Kata Containers (project site, read October 2026) | own machine, multi-tenant fleet |
| Managed remote sandbox | Someone else’s infrastructure; underneath, one of the three rows above, chosen by the provider | You cannot inspect the wall, only test it; session lifetime and stored state are the provider’s product decisions | E2B (documentation, read October 2026); Modal Sandboxes (documentation, read October 2026) | single-tenant service, multi-tenant fleet |
The last column records where those documents describe each category in use, and it ranks nothing. A missing tag means only that no document I read describes that use.
The book’s rule for choosing is a sentence it quotes from one vendor’s containment engineers: “Match isolation strength to the user’s capacity for oversight.” An engineer who reads commands fluently is part of the defense, so a lighter boundary with review is coherent on that engineer’s machine.
In a fleet, in Chapter 20’s words, strangers’ sessions run side by side on shared machines “and the wall between them is now the product”. No row removes the need for the other walls, and gVisor’s own security documentation says so: “A sandbox is not a substitute for a secure architecture.”
What do the book’s “battle-tested primitives” point at?
They point at three kinds of component and at no rung of the ladder. Chapter 17 advises building the boundary from “the container runtimes, syscall filters, and hypervisors that have absorbed decades of adversarial attention”, and it goes no further toward a choice.
The advice concerns the layer you add yourself. The chapter takes its reason from a vendor’s own account: “the weakest layer is the one you built yourself” (McGuinness and colleagues, 2026).
How does sandboxed code do authenticated work without standing access?
The real credential stays outside the sandbox and is attached to an outgoing request at a proxy, scoped to the task. The code holds at most a placeholder that is worthless anywhere else. Chapter 17 gives the logic as subtraction: “If the credentials file is never mounted inside the boundary, no injection can read it: there is nothing to outwit, because the target is absent.”
Chapter 20 turns that into a fleet default: “credentials injected per task and scoped to it, nothing standing”. Published designs reach it in two ways, each an example of the category:
- A proxy that checks, then signs. One coding agent’s sandbox keeps repository credentials outside and has a proxy verify each request before attaching the token (Dworken and Weller-Davies, 2025).
- A placeholder token. In one published design “the agent never sees real credentials. Instead, it receives a per-session authentication token that only works with a localhost proxy.” Its author notes that credentials which cannot pass through an HTTP proxy need another route (Hinds, 2026).
The containment write-up describes a desktop agent whose “Credentials stay in the host’s keychain and never enter the guest machine”, with “a per-session scoped-down token” that can be revoked separately from the user’s own.
Can AI agents be trusted with production access? For the code tool the question never has to be answered, because production credentials are absent from its room. An action that needs them goes through a named tool with its own scoped key and an approval gate.
What does an allow-list entry grant?
An allow-list entry grants every operation the destination offers, to every account the destination will act for. The containment write-up’s own lesson: “Every function reachable through any domain on an allowlist is now an attack surface.” Chapter 17 turns it into an instruction: “Enumerate what an allowed endpoint can be made to do before you call the list tight.” Hence the three parts of the policy’s network line: destination, operations, account.
How should results come back from the sandbox?
Results should come back as an exit code, a capped slice of output with any cut announced, and the paths of files written, while the bulk data stays inside. The model’s request is the easy half: structured output for tool calls makes the call parse, and it guarantees nothing about what the script prints.
Three rules from Chapter 5 cover what the script sends back:
- Cap the output and announce the cut. The chapter’s reason is that “a silent cut reads as a complete answer”.
- Return the exit code and the error text. The chapter tells you to “capture the errors and exit codes and feed them back”.
- Validate what came back. A clean exit is weak evidence: “that the program ran proves nothing about what it computed”.
Files are the fourth part. The workspace is inside the blast radius by construction, because the code has to write somewhere. The containment write-up says of its own virtual machine design that a compromised agent could still damage what is inside the workspace folder.
The blast radius of AI agents is set by three dials in Chapter 17, and the sandbox is one of them. For the workspace, the practical setting is a copy in and a reviewed diff out. One practitioner gave the reason: “there are all kinds of files that the agent could write to that I’d end up executing, as the developer, outside the sandbox” (Hacker News, March 2026).
A script stopped at the time limit will be run again, and any write it makes through a named tool has to survive the repeat. That is the subject of idempotent tools and safe retries.
How do you prove each wall holds?
You run a benign action from inside that the policy says must be refused or bounded, and you read the outcome in the operating system’s or the proxy’s own record. A configuration line, a status screen and the agent’s own statement are all claims. Anyone who says they sandbox AI agent tool execution should be able to show these records.
Each has misled someone. One issue on a coding agent’s tracker, from April 2026, reports a read denial that a status screen displayed while a combination of rules left it unenforced on Linux. The reporter found it enforced on macOS. I did not reproduce it, and it was closed as not planned.
The agent is the least reliable witness. One forum post describes an agent that said a file was outside its writable sandbox and, asked again, reported the edit done. The author traces it to a wrapper that grants agents all permissions by default, and comments: “the ‘pretend’ behaviour of giving up due to guardrails that are actually non-binding is pretty terrifying” (Hacker News, March 2026).
Before you start, plant two harmless markers: a canary file with a made-up token on the host, outside every mount, and a made-up credential in the environment of the process that launches the sandbox. Run the list as a script you wrote, in the deployment shape that serves your users. A test whose subject the policy sets to “none” is marked not applicable, with the reason.
- 1. Read a file outside the workspace. Open the canary file from inside. Expected: no such file, or a denial from the operating system where the boundary is a rule on the host’s own kernel. Seen in: the error text; in a boundary built on mounts, “access denied” means the file was mounted. Req. 1 · Filesystem: everything else
- 2. Change a read-only input. Write to an input the policy lists as read-only, and to a path outside the workspace. Expected: both refused by the operating system. Seen in: the host’s copies keep their checksums. Req. 1 · Filesystem: inputs, everything else
- 3. Compare the mounts with the policy. List the mounted paths from inside. Expected: the workspace and the named inputs only. Seen in: no home directory, credential folder, runtime control socket or policy file in the list. Req. 1 and 6 · Filesystem: never mounted
- 4. Look for a credential. Search the environment and the filesystem for the canary credential and for anything shaped like a key. Expected: nothing except the session’s placeholder, if the policy grants one. Seen in: the search output. Req. 3 · Credentials: inside the sandbox
- 5. Use the placeholder where it should be worthless. Present it to the destination directly, from outside the sandbox, then to the proxy after the session ends. Expected: rejected both times. Seen in: the destination’s error and the proxy’s log. Req. 3 · Credentials: real credential, revoked
- 6. Reach an address off the allow-list. Connect to a host the policy does not list. Expected: refused. Seen in: the egress log, with the session and tenant ids. Req. 2 · Network: default, logged
- 7. Ask an allowed destination for more than the policy grants. Request an operation the policy does not list, then a listed one as a second test account you own. Expected: both refused at the proxy. Seen in: two refused entries in the egress log. Req. 2 · Network: allowed
- 8. Exceed the time cap. Run a script that waits past the per-run limit, then keep a session past its per-session limit and leave one idle past the idle limit. Expected: the run stopped, both sessions destroyed. Seen in: a tool result that names the limit, and the teardown record. Req. 4 and 5 · Limits: wall-clock, on reaching a limit; Lifecycle: teardown
- 9. Exceed each resource cap. In staging, in separate runs, ask for more memory, processes, disk and processor time than allowed. Expected: each run stopped or held at its cap. Seen in: a second session beside it finishes normally. Req. 4 · Limits: processor, memory, processes, disk
- 10. Exceed the output cap. Print several times the limit. Expected: the result is cut at the limit and says how much was left out. Seen in: the tool result. Req. 4 and 7 · Limits: output; Results: truncation
- 11. Fail on purpose. Run a script that exits with an error. Expected: the exit code and error text return as a tool result, and the run continues. Seen in: the trace. Req. 7 · Results: returned
- 12. Leave a file for the next session. Write a marker, start a background process, end the session, start another. Expected: neither exists, and what the policy lists as surviving is in the outside store. Seen in: the new session’s file and process lists. Req. 5 · Lifecycle: teardown, must survive
- 13. See another session’s data. Run two sessions at once, as two test tenants if you have tenants, with a different marker in each. Expected: neither finds the other’s marker or processes, a connection from one session to the other is refused, and one tenant’s placeholder is rejected in the other’s session. Seen in: empty searches, and refused entries in the egress and proxy logs. Req. 2, 3 and 5 · Lifecycle: one sandbox per; Network: default; Credentials: real credential
- 14. Take the boundary away. In staging, make the sandbox unavailable and call the tool. Expected: an error, and no code runs anywhere. Seen in: the host’s process list during the call. Req. 8 · Boundary: if it cannot be set up
- 15. Follow a file out. Have the code write a file and end the session. Expected: it leaves only in the form the policy names. Seen in: nothing on the host has run or applied it before the stated check is recorded. Req. 9 · Results: files leaving
- 16. Repeat tests 1 and 6 through the agent. In a normal run, ask the agent to open the canary file and reach the unlisted host. Expected: the same refusals. Seen in: the operating system’s error and the egress log; the agent’s reply alone is no result. Req. 6 · Boundary: enforced by
What goes in the sandbox policy block?
The policy block holds one field for every wall the tests press on, plus two review items no test can press on: what the boundary is built from, and your own code on the trust path. It is plain text, so it belongs to no product. A policy to sandbox AI agent tool execution should read the same whatever enforces it, so translate this into your boundary’s configuration and keep it beside the code.
SANDBOX POLICY: [tool name] · owner: [team] · reviewed: [date]
Boundary (requirements 6 and 8)
Enforced by: [process restriction | container | user-space kernel |
microVM | managed service]
Built from: [named, maintained primitive]
Our own code on the trust path: [list, or none]
If the boundary cannot be set up: the tool refuses to run
Filesystem (requirement 1)
Workspace: [path], read-write, this session only
Inputs: [paths], read-only | none
Everything else: absent | denied by the kernel
Never mounted: home directory, key and credential folders,
the runtime's control socket, this policy
Network (requirement 2)
Default: deny
Allowed: [destination] -> [operations needed] -> [account] | none
Logged: every attempt, allowed or refused, with session and tenant ids
Credentials (requirement 3)
Inside the sandbox: none standing; [per-session placeholder | nothing]
Real credential: held [where, outside], attached at [proxy],
scoped to [task], expires [when] | none
Revoked: at session end, without touching the user's own
Limits, enforced outside the model (requirement 4)
Processor [ ] · memory [ ] · processes [ ] · disk [ ]
Wall-clock: per run [ ] · per session [ ]
Output returned to the model: [ ] characters
On reaching a limit: stop the run, return an error naming the limit
Results (requirements 7 and 9)
Returned: exit code, capped output and error text, paths of files
Truncation: announced in the result, with the amount left out
Files leaving: [a diff | named artifacts of listed types]
Before the host runs or applies them: [review by whom | type check]
Lifecycle (requirement 5)
One sandbox per: session and tenant
Teardown: files and processes destroyed at session end or after [ ] idle
Must survive: [what], written to [outside store] before teardown
Verified
Wall tests last run: [date] · by: [name] · in: [deployment shape]
Failed or not applicable: [test numbers, with reasons]
The caps are budgets in the book’s sense, and Chapter 20 describes them as budgets “enforced with walls instead of promises”.
What does the policy look like for a coding agent that runs tests?
It has a copy of the repository as the workspace, one registry on the allow-list for downloads only, and a diff as the only thing that leaves. I walked two designs through the table, the tests and the block to check that the three agree. This is the first, with illustrative numbers.
SANDBOX POLICY: run_tests · owner: developer tooling
Boundary microVM per session, from a maintained hypervisor · our own
code on the trust path: the egress proxy · no boundary, no run
Filesystem workspace: a copy of the repository, placed before the code
runs · input: the dependency cache, read-only · everything
else absent · never mounted: as in the template
Network deny · allowed: package registry -> download packages
-> one read-only account · logged with session and tenant ids
Credentials inside: a per-session placeholder · real: a read-only
registry token, held in the secret store, attached at the
proxy, scoped to downloads, expires and is revoked at session end
Limits 2 cores · 4 GB · 256 processes · 10 GB disk · 10 min per run
· 60 min per session · 20,000 characters returned · at a
limit: run stopped, error names the limit
Results exit code, capped output, file paths · truncation announced
· leaving: a diff and the test report · before the host
applies them: review by the requesting engineer
Lifecycle one per session and tenant · destroyed at end or after
10 min idle · must survive: the diff and the report, written
to the artifact store first
Verified 2026-10-06, by the owner, staging pool on the production
image · failed or not applicable: none
All sixteen tests apply. Test 7 matters most here: a registry accepts uploads as well as downloads, so the proxy has to refuse a publish request and any request made as another account. No repository credential appears, because the copy goes in before the code runs and the change comes out as a diff.
What changes for a data-analysis agent?
Five lines change and two tests drop out. The second design runs model-written code on an uploaded file and returns a chart, for many customers. Its input is the upload, mounted read-only, and its tenant line names each customer. Its allow-list is “none” and its credentials line is “nothing”, because the image carries the libraries it needs.
Tests 5 and 7 become not applicable, as does the placeholder part of test 13, and test 6 carries the whole network wall. The fifth line is the chart, which leaves as a named artifact of one listed image type, checked by type before the product displays it. Every field could be filled for both designs, and no test asked for something the policy forbids.
What does a sandbox not stop?
A sandbox does not stop a wrong program, misuse of a path you allowed, damage to what you mounted, or the injection itself. Lists of AI agent security best practices often end at “use a sandbox”. The published incidents sort into two groups.
With a sandbox in place. One vendor’s containment write-up describes a case, which it says came from a third-party disclosure, of a desktop agent running in a virtual machine with an egress allow-list. A malicious file in the user’s workspace carried instructions and an API key that belonged to the attacker. The agent uploaded workspace files through the vendor’s own allow-listed API domain, into the attacker’s account.
The write-up’s verdict: “The sandbox worked perfectly, and yet the data was exfiltrated” (McGuinness and colleagues, 25 May 2026). That is evidence for the other rings: the review of each allow-list entry, and test 7.
With no wall that stopped it. The same write-up reports a controlled red-team exercise from February 2026, and does not say which boundary, if any, was switched on. An employee was phished into launching a coding agent with a prepared prompt, and the agent completed the exfiltration in 24 of 25 retries. Chapter 17’s reading is that “The only layer that could have stopped it was the environment”.
The database loss from the opening belongs here too. The author’s stated changes afterward include “Agents no longer execute commands” and “Every destructive action is run by me” (Grigorev, 6 March 2026). Those are the approval dial, and the write-up never mentions a sandbox or credentials. What it shows for this post is row 3: commands the agent ran had standing access that could destroy production.
A third report shows why the agent’s account of the filesystem is weak evidence: in a July 2025 issue on a coding command-line tool, one user reports an agent that reorganized a folder and declared the files lost, and the client details list no sandbox. A week later the reporter wrote that the files had turned up at the root of the drive, outside the project folder, and that the bug was “not nearly as severe”. No data was lost. With no boundary, a file move landed outside the folder the agent was working in, a write that test 2 would refuse.
Two more things stay outside the room’s reach. The first is the wrong program, in Chapter 5’s words: “The model can still write the wrong program, which will then produce the wrong answer with perfect consistency”. The second is the injection rate; Chapter 17 says of its three dials that “All three leave the injection rate alone”. The neighboring post is the lethal trifecta explained.
What the room buys beyond containment is freedom inside it. Chapter 17 observes that “a tight perimeter is what makes it safe to relax oversight inside”. By one vendor’s account an operating-system sandbox brought “an 84% reduction in permission prompts” (McGuinness and colleagues, 2026). Approval then goes to what leaves the room, and the guide to AI agent guardrails covers how those gates are built.
The lethal trifecta audit scores the sandbox dial with four practices that match rows 2, 3 and 6 of the table. The consequence tier classifier sorts the actions that leave the sandbox into the ones that wait for a person.
What changes at one sandbox per session and at scale?
Three things change: the wall starts protecting tenants from each other, the room has to exist before the user asks for it, and teardown becomes a security control. Chapter 20 frames the shift: “One sandbox is a security decision. Ten thousand a day is an infrastructure product, with its own economics of boot time, pooling, and teardown”. The requirements to sandbox AI agent tool execution stay the same nine, and the defaults harden.
| Question | One sandbox on your machine | One per session, in a fleet |
|---|---|---|
| Who does the wall protect? | You, from your own agent’s code | Each tenant from every other tenant |
| When does the room start? | When you ask | Ahead of demand, in a warm pool |
| What survives a run? | Whatever was left behind | Only what the policy wrote outside before teardown |
| Where are the caps? | A timeout somebody remembered | Processor, memory, wall-clock and output size, per room |
| Where do credentials come from? | Often your own environment | Injected per task, scoped to it, nothing standing |
| Where does the code run? | In the agent’s process | In a pool of executors scaled on its own |
| What does the bill show? | Nothing you notice | Uptime, a cost outside the token count |
The warm pool exists because of a trade, and Chapter 20 states it: “Strong isolation costs time at boot, and a user is often waiting, so fleets keep a warm pool: sandboxes booted ahead of demand, claimed in milliseconds, one per session.” The chapter reads the ladder of isolation strengths, at scale, “as a single trade of boot time against wall thickness”. I print no boot times, because every figure I found was a vendor’s own or lacked stated conditions. One forum commenter asked what a boot faster than the model’s first token buys in an agent loop (Hacker News, March 2026). I have no answer that favors the faster boot.
One company’s fleet write-up from October 2026 shows the same defaults in practice: “We keep a pool of machines booted ahead of demand”. It keeps a session’s machine alive for a set interval after each reply, the policy’s idle line. It also states the rule behind row 8: “Fail closed. If any of this can’t be set up, the session doesn’t start.” (Pulavarthi, 2026)
On teardown, Chapter 20 is blunt: “anything that survives teardown is a message from one session to the next”. The chapter names the new hazard in one line: “The hazard the demo never had is other people.” Tests 12 and 13 are the two to run on every release of the pool.
The table’s last row is this post’s addition, and the others are Chapter 20’s. One practitioner’s thread asks how to track agent costs that are not tokens, sandbox time among them (r/LLMDevs, September 2026).
Where does this guide stop?
It stops at evidence that a wall held for the actions you tried, on one day, in one deployment shape. Sixteen passing tests are silent on a determined adversary and on flaws in the primitive.
A managed provider’s wall is closed to inspection, so the tests are the only evidence you get.
Browser and computer-use tools raise their own questions, which Chapter 5 takes up in its next section. Whether to have a code tool at all is row 4 of the AI agent architecture guide.
Test the wall you already have
To sandbox AI agent tool execution is to make nine claims about a room, and each claim either produces a refusal you can read in a log or it does not. Fill in the block, run the sixteen tests, and file what fails.
Chapter 2 is free to read and explains why exact work goes to something exact, and the glossary is free too. Chapter 5, “Tools and the Action Space,” Chapter 17, “Security, Safety, and Guardrails,” and Chapter 20, “Deploying and Scaling,” are in the full book. The tools and protocols guide collects the neighboring posts, or you can see the formats.
Questions readers ask
- Is a container enough to sandbox an AI agent's code?
- It can be for one engineer on their own machine who reads what runs, with nothing sensitive mounted and the network denied. A container shares the host's kernel, and one widely used runtime's security documentation says its defaults may give incomplete isolation. For strangers' sessions on shared machines, Chapter 20 of AI Agents, Engineered sets stricter defaults: one sandbox per session and per tenant, egress denied by default, credentials injected per task.
- How do I let sandboxed code call an API without giving it the key?
- Keep the real credential outside the sandbox and attach it to the outgoing request at a proxy, scoped to the task and the session. The code holds at most a placeholder that is worthless anywhere else. Published examples of the pattern include a proxy that checks the request before adding the token, and a per-session placeholder token that only a local proxy accepts (Hinds, 2026).
- Does a sandbox stop data exfiltration?
- It stops the routes you closed. One vendor's 2026 containment write-up describes workspace files leaving through the vendor's own allow-listed API domain, into an attacker's account, with the sandbox working as designed. For every allowed destination, list the operations it offers and the accounts it will act for, then test that anything outside that list is refused.
- How do I know my sandbox is working?
- Run benign actions from inside that the policy says must fail: open a file outside the workspace, connect to a host off the allow-list, look for a planted credential, exceed the time and output caps, leave a file for the next session. Read each refusal in the operating system's or the proxy's log. A status screen, a configuration line and the agent's own statement are all claims to be tested.
- What should a code execution tool return to the model?
- The exit code, a capped slice of standard output and standard error with any truncation announced in the result, and the paths of files the code wrote. The bulk data stays inside the sandbox. Chapter 5 of AI Agents, Engineered gives the reason for announcing the cut: a silent cut reads as a complete answer.
Sources
- McGuinness, Grace, De Jonghe, Eaton, Ribbink (Anthropic) (2026). How we contain Claude across products
- Dworken, Weller-Davies (Anthropic) (2025). Beyond permission prompts: making Claude Code more secure and autonomous
- Varda, Pai (Cloudflare) (2025). Code Mode: the better way to use MCP
- Alexey Grigorev (2026). How I Dropped Our Production Database and Now Pay 10% More for AWS
- Atchyut Pulavarthi (QA Wolf) (2026). We gave every AI agent its own computer after unlearning cloud development
- Luke Hinds (2026). Credential Protection for AI Agents: The Phantom Token Pattern
- Agache, Brooker, Florescu, Iordache, Liguori, Neugebauer, Piwonka, Popa (2020). Firecracker: Lightweight Virtualization for Serverless Applications (NSDI '20)
- Souppaya, Morello, Scarfone (2017). Application Container Security Guide (NIST Special Publication 800-190)
- The Linux Kernel documentation (2026). Landlock: unprivileged access control (kernel documentation, page dated August 2026)
- bubblewrap project (2026). bubblewrap README (read 2026-10-06)
- Docker (2026). Docker Engine security (documentation, read 2026-10-06)
- Docker (2026). Docker Sandboxes: Architecture (documentation, read 2026-10-06)
- gVisor project (2026). What is gVisor? (documentation, read 2026-10-06)
- gVisor project (2026). gVisor Security Model (documentation, read 2026-10-06)
- Kata Containers project (2026). Kata Containers project site (read 2026-10-06)
- E2B (2026). E2B documentation (read 2026-10-06)
- Modal (2026). Sandboxes (Modal documentation, read 2026-10-06)
- anuraag2601 (GitHub) (2025). Gemini CLI "lost" files during a failed file move operation (GitHub issue #4586, closed as not planned; the reporter later found the files)
- riywo (GitHub) (2026). A read denial that is displayed and not enforced (GitHub issue #53209, closed as not planned)
- ekorondy (Hacker News) (2026). Hacker News comment asking for an honest matrix of what each sandbox layer stops
- benjosaur (Hacker News) (2026). Tell HN: AI lies about having sandbox guardrails
- rsyring (Hacker News) (2026). Agent sandboxing tools that mount projects R/W have gaping exploits (Hacker News)
- ammmir (Hacker News) (2026). Hacker News comment on cold starts and the model's time to first token
- u/Any_Warning_1183 (Reddit) (2026). How do you track agent costs that aren't tokens? (r/LLMDevs)