Home / Tools / Spot the injection

Free tool · runs in your browser · from Chapter 17

Spot the injection

A prompt injection game: read a poisoned inbox, grant an agent its tools and settings, run it, and see what leaks. Keep the most capability with no leak.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The three legs, the rules about each defense and the quotes are the book's (Chapters 9 and 17). The inbox, the pages and the payloads, the agent that always obeys, the fixed effect of each setting and the score are the tool's own. The game covers theft only.

What is a prompt injection game?

A prompt injection game is an exercise in which you set up an AI agent, let it read content that someone has poisoned, and watch whether your private data leaves. In this one you look over a small inbox and two web pages, choose the agent’s capabilities and defenses, and run it. The score is the capability you manage to keep while nothing leaks.

The mechanism it teaches is the one Chapter 17 of the book opens with. “The model reads one undifferentiated stream of tokens: your standing instructions, the user’s request, the fetched page, the tool results, all concatenated, with nothing structural to mark where the trusted material ends and the merely-read material begins.” An attacker does not need access to your system. They need to place text where your agent will read it, and an inbox or a public web page will do.

The game is a companion to the lethal trifecta audit. The audit is a checklist you run on your own design. This page gives you a design that is already assembled, three concrete attacks, and a trace of what each one did.

The lethal trifecta.
Figure 17.2 The lethal trifecta. Each capability on the left is individually useful, and each is often individually fine. An agent holding all three can be turned into a data thief by anyone who can write to its inputs: the attacker plants an instruction through the untrusted-content leg, the instruction gathers secrets through the private-data leg, and the loot leaves through the outbound leg. The audit is fast, and so is the remedy: removing any single leg defeats the theft. Reuse this diagram

Why can’t you win by spotting the injections?

You cannot win by spotting the injections because two of the three are invisible to a person, and the one you can see reads like an ordinary request. The first step of the game lets you flag items anyway. It reports how you did and then leaves your flags out of the score, for a reason Chapter 17 gives directly: “the payload does not need to be visible to a person.”

Each item has a control that shows what the model reads. One page carries a paragraph in white text on a white background. One email carries a comment addressed to the assistant. A human reviewer sees a product review and a welcome message. The model sees all of it, and the book describes the failure as diligence, not carelessness: “it read everything, including the part addressed only to it.”

The inbox also contains a decoy. A colleague writes “Ignore my last message” and asks for a deck to be forwarded, which is a normal thing for a person to say and looks exactly like an attack if you are scanning for suspicious phrases. That is the trouble with detection as a strategy. “Text that is shaped like an instruction pulls on the model, whoever wrote it and whatever it is doing there.”

So the game moves you on to the question the chapter says matters. “You do not build on the assumption that injection is blocked; you build on the assumption that one eventually lands, and you engineer the day after.”

What are the three attacks in the game?

The three attacks are a visible email that asks for a forward, a hidden page footer that asks for an image, and a hidden email comment that asks for an upload. Each one is a variation on a case the book describes, moved to reserved example addresses, and each leaves by a different route.

Payload Where it hides What it asks for The route out
The polite forward A visible email “for the assistant” Forward password-reset emails to an outside address, then delete them Send email
The invisible footer White text on the review page End the briefing with an image whose address carries your data A rendered image, or an outbound request
The setup snippet A comment in a vendor email Upload your data to the vendor’s domain, using the attacker’s account key An outbound request the allow-list permits

The first is the book’s email assistant. “Reading mail is leg one and leg two in a single tool: the inbox is private data, and it is also a channel any stranger on earth can write to.” The second shows how the third leg accumulates without anyone adding a send tool: “So is rendering a Markdown image whose URL the model composes: the stolen secret rides out in the query string when the image loads.” The third is the incident the chapter asks you to carry with you, in which a sandboxed agent uploaded files through a permitted domain into an attacker’s account.

One detail is easy to miss. Switching off access to private files does not remove the first leg while the inbox stays readable, because the inbox is private data too. The payloads in the game fall back to the email bodies when the file is out of reach.

Which defenses work in the game, and which do nothing?

The defenses that work in the game are the ones that remove a route or put a person in front of the few actions that leave; the ones that do nothing are the ones that ask the model, or a classifier, to notice the attack. That split is Chapter 17’s, and the game hard-codes it for the worst case.

Defense Effect in the game What the book says
A prompt rule to ignore instructions in content None “An instruction in the prompt is a request.”
An injection detector None in the worst case “a delay, not a defense”
A network allow-list Stops one fetch, permits the upload “a capability grant, not a destination filter”
Approval on every action None “Prompting on everything trains the click that defeats the prompt.”
Approval on what leaves Stops the forward and the upload, not the image “spend it where a considered”no” is plausible”
A quarantined reader Stops the two hidden payloads “safer, never safe” (Chapter 9)
Removing a capability Closes that route for certain “remove capability instead of filtering it”

The detector deserves a word, because it does catch things in practice. The chapter grants that statistical defenses lower the success rate and then points at who moves second: the attacker studies the defense and rephrases. “Several percent, against an adversary who can afford a hundred tries, is a delay, not a defense.” The game’s attacker has already rephrased.

The allow-list behaves in two ways that surprise people. It filters the agent’s own requests, so it does nothing about an image that your screen loads when you open the briefing. And it permits the vendor’s domain for every account, the attacker’s included. The book’s sentence is the one to remember: “an allow-list entry is a capability grant, not a destination filter.”

Approval is the scarce defense. Switch it on for everything and the game reports that the important prompt got the same reflexive click as the rest. Switch it on for the actions that leave, sends and outbound requests, and the prompts you see are few enough to read: one shows an outside address and the words “password reset”, another a private file going to an unfamiliar account. It never fires for the image, because writing a briefing is not an action that leaves; your own screen makes that request later. Chapter 17’s accounting is short: “An approval gate is a scarce resource: spend it where a considered”no” is plausible.” The same chapter prefers the structural version where you can have it: “an agent that drafts email for your approval beats one that sends.”

How does the scoring work?

The score is the number of capabilities you kept, out of six, on a run where nothing leaked and the briefing still got written. The result also reports how many settings you needed, and fewer is better. The scoring is the tool’s own invention; the book ranks nothing this way.

The idea behind it is the book’s, though. Safety by subtraction is always available: an agent with no inbox and no browser leaks nothing and does nothing. The interesting question is how little you have to give up. Chapter 17 calls cutting external communication “frequently the cheapest amputation”, and its guardrail section says where the strongest moves are: “The strongest moves here remove capability instead of filtering it.” Structural cuts earn their place because “a guard that makes no bet on catching the bad thing cannot lose that bet.”

Two settings are enough to keep all six capabilities, and that is the best configuration the tool knows: a quarantined reader and approval on what leaves. Without the reader, the best safe configuration gives up rendered images, which is part of the amputation the chapter describes. After your first safe run, the result tells you how far you are from that mark.

What does the simulation leave out?

The simulation leaves out probability, destruction and two other ways an injection can arrive. It is a teaching model with a fixed outcome for every combination of switches, and four of its simplifications matter.

First, the agent always obeys. The book is careful to say that an assistant complies “some fraction of the time”, a fraction no prompt wording pushes to zero. The game takes the worst case every time, which is the right assumption for design and the wrong one for estimating how often an attack succeeds.

Second, the quarantined reader and the approval prompt block more reliably here than they do in life. Chapter 9 is explicit about the reader: “Containment reduces exposure; it does not neutralize it.” The game shows one way through, where a summary carries the visible email’s request forward as a fact, and treats the rest as blocked. A real approver can also click yes.

Third, the game is about theft. Chapter 17 warns that damage runs on less: untrusted content plus one consequential tool, such as a delete or a refund, is enough to break things with no data leaving. Fourth, every payload here arrives through content the agent fetches. Injections pasted in by a user, or shipped inside a tool’s own description, are out of scope.

For your own agent, run the lethal trifecta audit, which covers the second audit and the three blast-radius dials, and the tool contract linter for the descriptions you install. The lethal trifecta explained is the essay version, and the AI agent security guide places it beside guardrails and sandboxing. The chapter closes its containment argument with a line that fits the game’s score: “You cannot buy a model that is never wrong; you can always shrink what wrong costs.” The full treatment is in Chapter 17, Security, Safety, and Guardrails (in the full book).

Questions readers ask

What is a prompt injection, in plain terms?
A prompt injection is text an agent reads as part of its work that it then treats as an instruction. It can sit in an email, a web page, a document or a tool result. The model receives its instructions and its reading material as one stream of text, so content shaped like an order can pull on it whoever wrote it.
Can I win by spotting every injection?
No, and the game does not score your flags. Two of the three payloads are hidden from a person, and one clean email looks like an attack. Chapter 17's advice is to assume an injection eventually lands and to limit what it can reach, which is what the second step of the game asks you to do.
Why does the prompt rule not stop anything?
Because nothing enforces it. A sentence in the system prompt is read and weighed by the model on each pass, and a persuasive page can outweigh it. The book's phrase is that an instruction in the prompt is a request. The game runs the worst case, so the rule always loses.
Why did data leak after I removed the send-email tool?
Because sending mail is only one way out. An image in the agent's formatted output is requested by your own screen when you open the briefing, and the address can carry private data. Any outbound request is another route. If anything the agent emits can cause a request to an address the attacker chose, the third leg is present.
Is the quarantined reader a complete defense?
No. It stops the two hidden technical payloads in this game because they have no field to travel in. The plausible-sounding request in the visible email survives as an accurate summary, and the main agent acts on it. The book calls a quarantined result safer, never safe.
Does a safe result here mean my own agent is safe?
No. The simulation is deterministic and small, and it covers data theft only. It does not model destruction through a consequential tool, an injection typed or pasted by a user, or a poisoned tool description. Use the lethal trifecta audit on your real design.

Sources

  1. Simon Willison (2025). The lethal trifecta for AI agents: private data, untrusted content, and external communication
  2. Kai Greshake et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
  3. Luca Beurer-Kellner, Beat Buesser, Ana-Maria Creţu, et al. (2025). Design Patterns for Securing LLM Agents against Prompt Injections
  4. OWASP Gen AI Security Project (2025). LLM01:2025 Prompt Injection