Learn / Explainers / The lethal trifecta: three legs, one theft, and the cut

The lethal trifecta: three legs, one theft, and the cut

Animated explainer · 3:45 · from Chapter 17

Space plays and pauses · ← → step · F fullscreen

A short animated explainer of the lethal trifecta: three agent capabilities that are fine alone, the theft when they meet, and the one leg to cut.

Transcript

Every step of the animation, in the words it uses on screen. Select a line to jump to it.

From Chapter 17

Security, Safety, and Guardrails

This explainer condenses a few pages of Chapter 17, which is in the full book. The chapter traces the ideas through a worked example; the preface and Part I are free to read online.

Questions

Who coined the term "lethal trifecta"?
Simon Willison, in a June 2025 post titled “The lethal trifecta for AI agents: private data, untrusted content, and external communication.” Chapter 17 adopts his framing as its audit: count which of the three capabilities an agent holds, and treat all three together as a standing invitation to theft. The lethal trifecta explained post walks through the framing and its incidents in more depth.
What is the cheapest leg to remove?
Chapter 17 goes leg by leg in this order. First, can the agent run without private data? Then the risk collapses to vandalism. Second, can it run on a closed world of inputs, with nothing attacker-writable in its diet? Rare in practice. Third, can you cut external communication? The chapter calls this “frequently the cheapest amputation”: deny network egress by default, allow a short list of destinations, and disable link rendering and image loading in the agent’s outputs. Run the lethal trifecta audit to see which legs your own design holds.
Isn't prompt injection the model vendor's problem to fix?
No. Jailbreaking targets a model’s safety training, and the embarrassment mostly lands on the vendor; prompt injection targets your application, and the tools, credentials, and data it reaches are yours. Every defense at the model layer is statistical, so build on the assumption that an injection eventually lands and cap what it can cost. The agent security and operations guide covers the containment around the agent: deterministic guardrails, least privilege, sandboxing, and human approval.
If the agent holds only two legs, is it safe?
Safe from theft, not from damage. The trifecta describes exfiltration. An agent holding untrusted content plus any consequential tool (delete, refund, merge, deploy) can be goaded into destruction with no private data leaving anywhere. So run both audits: three legs mean an attacker can steal; two legs and a sharp tool mean an attacker can break.