Home / Glossary / prompt injection

prompt injection

The attack family in which instructions are smuggled into text the model reads, exploiting the fact that the model sees one undifferentiated token stream with no structural boundary between trusted instructions and merely-read content. Direct: the attacker is the user, typing. Indirect: the instruction hides in content the agent encounters doing legitimate work—a web page, an email, a code comment—and this is the main event for tool-using agents, whose job is reading things other people wrote. Every known defense is statistical; this book's working rule is to treat any claim of prevention as false and to engineer the blast radius instead (Chapter 17).