Home / Tools / Lethal trifecta audit

Free tool · runs in your browser · from Chapter 17

Lethal trifecta audit

A lethal trifecta checklist for AI agents: see which legs your design holds, whether a sharp tool adds risk, and the cheapest leg to cut. Print the audit.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The three legs, the second audit, the verdicts, the mitigations and the dial practices are Chapter 17's. The grouping of the checklist items, the order of the verdicts and the dial scores (a count of practices in place) are the tool's own.

What is a lethal trifecta checklist?

A lethal trifecta checklist is a short audit of three capabilities an AI agent may hold: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent with all three can be turned into a data thief by anyone who can write to its inputs. Removing any one leg defeats the theft.

The framing comes from Simon Willison and is the center of Chapter 17 of the book, which calls it “the single most useful piece of security thinking in this book.” The instruction is plain: “Count which of three capabilities your agent holds.” The tool above turns that count into a checklist you can run in a few minutes, adds the second audit the chapter insists on, and scores the three dials that decide how much a successful attack costs. The lethal trifecta explained goes through the idea at essay length; this page is the working reference.

The lethal trifecta.
Figure 17.2 The lethal trifecta. Each capability on the left is individually useful, and each is often individually fine. An agent holding all three can be turned into a data thief by anyone who can write to its inputs: the attacker plants an instruction through the untrusted-content leg, the instruction gathers secrets through the private-data leg, and the loot leaves through the outbound leg. The audit is fast, and so is the remedy: removing any single leg defeats the theft. Reuse this diagram

What are the three legs?

The three legs are access to private data, exposure to untrusted content and the ability to communicate externally: three capabilities that are each useful and often harmless on their own, and that together let anyone who can write to the agent’s inputs steer it into theft. Chapter 17 says so directly: “Each capability is individually useful and often individually fine.” The danger is the combination, because “an attacker plants an instruction through leg two, the instruction directs the agent to gather secrets through leg one and ship them through leg three.”

Leg The book’s definition What the checklist asks about Cheapest way to remove it
1. Private data “your files, your mail, your customer records” Files and repositories, messages, records, reachable credentials Use only data you would be content to see published
2. Untrusted content “any channel by which text or images an attacker controls can reach the model’s context” Web pages, inboxes and issue trackers, outside documents, third-party tool results, subagent summaries, installed tool descriptions, pasted prompts A closed world of inputs, or a quarantined reader
3. External communication “any way to move information out” Send tools, any HTTP request, rendered images and links, allow-listed domains Deny egress by default; disable link and image rendering

Leg two is wider than most teams assume. A subagent’s tidy summary is the output of whatever it read, which Chapter 11 of the book describes as an injected instruction “laundered into a voice your system was built to trust.” A tool description installed from a marketplace is read with the same obedience as anything else in the context, which is how tool poisoning works. And a prompt a colleague pastes in from a message counts too: Chapter 17 reports a red-team exercise in which a phished employee pasted a ready-made prompt, and the agent completed the exfiltration in 24 of 25 retries.

Why is the third leg so easy to miss?

The third leg is easy to miss because external communication rarely looks like a send button: an outbound request, a rendered image or a clickable link can carry a secret to an address the attacker chooses, even when the agent’s tool list contains nothing named send. Engineers picture a send_email tool they carefully left out. The chapter’s correction is the paragraph to remember: “Any HTTP request is egress. So is rendering a Markdown image whose URL the model composes: the stolen secret rides out in the query string when the image loads. So is producing a hyperlink the user might click.”

The test that follows is mechanical. If anything the agent emits can cause a network request to an address an attacker influences, leg three is present, whatever the tool list says. That is why the checklist asks about rendering and allow-lists, not only about tools.

Allow-lists deserve their own line because they feel safe. The vendor postmortem Chapter 17 cites describes a sandboxed agent whose network access was limited to the vendor’s own API domain. A malicious file supplied the attacker’s API key, and the agent uploaded workspace files through the permitted domain into the attacker’s account. “The sandbox worked perfectly, and yet the data was exfiltrated.” The book’s rule: “an allow-list entry is a capability grant, not a destination filter.”

How does the audit work on the book’s examples?

The audit works by ticking every capability the agent can actually reach, then reading the verdict: on the book’s email assistant, three ticks (the mail it reads, an inbox strangers can write to, a send tool) give theft, and the cheapest cut is external communication. Two buttons above load the book’s own examples, and both come out at three legs.

The first is the email assistant, the agent most teams were planning to build anyway. “Reading mail is leg one and leg two in a single tool: the inbox is private data, and it is also a channel any stranger on earth can write to. Sending mail is leg three.” The checklist needs three ticks: messages, an inbox strangers can write to, a send tool. The verdict is theft, and the cheapest amputation is leg three: drafts for approval instead of sends, no rendered images or links, egress denied by default.

The second is the code-host connector incident from Invariant Labs’ 2025 disclosure. A connector could read issues on public repositories, read private repositories, and open pull requests. An attacker filed an issue carrying instructions; the owner asked the agent to look at open issues; the agent copied private contents into a pull request on the public repository. The researchers’ conclusion, as Chapter 17 relays it, is that “there was no bug in the connector’s code.” The vulnerability was the combination, and only whoever assembled the combination can fix it. Increasingly, the book notes, that assembler “is you, on a Tuesday, adding one more connector because it looked useful.”

Why run a second audit for sharp tools?

You run a second audit because the trifecta describes only theft, and damage needs less: an agent that reads untrusted content and holds even one consequential tool, such as delete, refund, merge or deploy, can be steered into real harm without any private data leaving. Chapter 17 draws the line itself: “An agent holding untrusted content plus any consequential tool—delete, refund, merge, deploy—can be goaded into destruction with no private data leaving anywhere: two legs, no exfiltration, real harm.”

The condition that makes this possible has a name in the security community, excessive agency: more permissions, functionality or freedom than the task requires. OWASP lists it among its top risks for applications built on large language models. The book’s summary is the sentence the tool’s verdict is built on: “Three legs mean an attacker can steal; two legs and a sharp tool mean an attacker can break.”

The remedy for a sharp tool is to remove capability rather than filter it. An agent whose only path to money is a typed refund(order_id, amount) with the cap enforced inside the function has nothing to inspect, because the dangerous capability was never granted. Anything irreversible that remains goes behind an approval gate.

Which leg should you cut first?

You should usually cut leg three first, because the book calls it “frequently the cheapest amputation.” The chapter walks the legs in order, and the tool reports the same walk for every leg your agent holds:

  1. Private data. Can the agent run without it, “or with access only to data you would be content to see published? The risk collapses to vandalism.”
  2. Untrusted content. Can it run “on a closed world of inputs, nothing attacker-writable in its diet? Rare in practice.” Where it cannot, a quarantine pattern helps: a privileged agent that never reads raw untrusted text delegates reading to a worker whose findings come back as inert, structured data.
  3. External communication. Cut it: “deny network egress by default, allow a short list of destinations, disable link rendering and image loading in the agent’s outputs, and leg three withers.”

The research paper the chapter quotes states the target condition: “once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions.”

What are the three blast-radius dials?

The three blast-radius dials are the settings that cap what a successful injection can do: “what the agent is permitted, where it runs, and which of its actions wait for a human.” The book calls them least privilege, the sandbox and human approval, and its point is that all three work on the price of an attack, not its rate. The tool scores each dial by counting the book’s practices you already have in place.

Least privilege cuts the deputy’s keys per errand: read-only where reading suffices, a credential scoped to one project, a tool that queries one table, drafts instead of sends. The sandbox removes targets: if the credentials file is never mounted inside the boundary, no injection can read it. Build it from proven container runtimes, syscall filters and hypervisors, because “the weakest layer is the one you built yourself.” Approval is the scarce dial. According to Anthropic’s 2026 containment write-up, users approved roughly 93% of permission prompts, and the book draws the consequence: “Prompting on everything trains the click that defeats the prompt.”

The tool also asks who is watching. The rule comes from the same disclosures: “Match isolation strength to the user’s capacity for oversight.” An engineer who reads shell commands can be part of the defense. A non-technical user cannot, so the tool stops counting approval as a defense for that operator, and for unattended runs, where no one is there to click.

What does the audit not tell you?

The audit does not tell you how likely an attack is, and it cannot make prompt injection impossible; it tells you which legs make theft possible, which tools make destruction possible and what a successful injection would be worth, so you can lower that price. The book’s working rule is blunt: “any claim that a prompt, a product, or a model prevents prompt injection should be treated as false.” Willison makes the same point about detection products that advertise catching 95% of attacks: in web application security, 95% “is very much a failing grade.”

The checklist is also only as good as your ticks. It cannot see a connector you forgot, and a dial score counts practices, not how well each one is built. Use it as the one-minute check before a design review, then re-run it every time a tool or connector is added. The AI agent security guide puts the audit next to guardrails, reliability and cost. When you know how expensive a wrong action is, the consequence tier classifier helps decide which actions need a gate, and should this be an agent? asks the earlier question of whether the design needs this much autonomy at all.

Chapter 17 closes the containment argument with a line worth keeping in view while you audit: “The deterministic boundary is what gets hit when everything probabilistic misses.” The full treatment, including guardrails and the supply chain, is in Chapter 17, Security, Safety, and Guardrails (in the full book); the related posts on how to prevent prompt injection, indirect prompt injection attacks and the blast radius of AI agents take each piece further.

Questions readers ask

My agent has no send_email tool. How can leg three be present?
Because egress rarely looks like a send tool. Any HTTP request is egress, and so is a Markdown image whose URL the model composes, since the secret leaves in the query string when the image loads. A link the user might click counts too. If anything the agent emits can reach an address an attacker influences, leg three is present.
Why does the audit say removing one leg is enough?
Because the theft needs all three. The attacker plants an instruction through untrusted content, the agent gathers secrets through private data, and the loot leaves through external communication. Take away any one and the chain breaks. That is a statement about theft only; destruction needs less, which is why the tool runs a second audit.
What is the second audit?
Untrusted content plus any consequential tool, such as delete, refund, merge or deploy. An injected instruction can make the agent break things with no data leaving at all. Chapter 17 puts it in one line: three legs mean an attacker can steal; two legs and a sharp tool mean an attacker can break.
Why does the book say approval prompts can defeat themselves?
Because people stop reading them. One vendor's telemetry found users approving roughly 93% of permission prompts, and the more prompts people see, the less attention each gets. Prompting on everything trains the reflex that defeats the prompt, so the book reserves approval for the few consequential, hard-to-reverse actions.
Does passing this audit mean my agent is safe from prompt injection?
No. Nothing in the audit lowers the rate at which injections succeed; it lowers what a successful one is worth. The book's standing rule is that any claim that a prompt, a product or a model prevents prompt injection should be treated as false.

Sources

  1. Simon Willison (2025). The lethal trifecta for AI agents: private data, untrusted content, and external communication
  2. Anthropic (2026). How we contain Claude across products
  3. Invariant Labs (2025). GitHub MCP Exploited: Accessing private repositories via MCP
  4. Beurer-Kellner et al. (2025). Design Patterns for Securing LLM Agents against Prompt Injections
  5. OWASP Gen AI Security Project (2025). LLM06:2025 Excessive Agency