Home / Tools / Multi-agent failure explorer

Free tool · runs in your browser · from Chapter 11

Multi-agent failure explorer

Pick the symptom your multi-agent run showed, read the seam behind it and the fix, check a worker brief for four fields, and decide whether to split.

The tool

Your inputs stay in this tab. Share a result by copying the page address: the state lives in the URL. The three clusters, the four seams and their examples, compounding across agents, the single-threaded-writes rule, the four fields of a brief and the contractor test, the merge wall and its remedies, and the question that decides a split are the book's (Chapter 11). The tool's own: the list of symptoms and the card each opens, the word patterns of the brief check, the 12-word flag, the wording of the questions and the rule that turns answers into a verdict.

Why do multi-agent systems fail?

Multi-agent systems fail mostly at the points where one agent hands something to another. Chapter 11 of the book calls those points seams: “With five agents you own a distributed system, and its bugs live in the seams: a dropped handoff, an assumption two workers silently made differently, a summary that shaved off the one detail that mattered.”

The chapter leans on a published study of execution traces. A research group annotated more than 1,600 traces from seven multi-agent frameworks and sorted the recurring failures into “fourteen modes in three clusters: system design issues, inter-agent misalignment, and weak task verification”. The book names the clusters and does not list the fourteen modes, and this page does the same. It covers what the chapter describes in its own words: four seams, the arithmetic underneath them, the merge wall, and two operating facts about cost and repeatability.

The tool above has three parts, and each works alone. The first starts from what you saw go wrong. The second reads a worker’s brief. The third asks whether the work should have been split in the first place.

How do I find which seam failed?

You find the seam by starting from the symptom and reading back to the crossing where it was introduced. Tick every symptom that matches the run, and the tool opens the card for each. The list of nine symptoms and the card each one opens are the tool’s own. Every card restates Chapter 11 and quotes it.

What you saw Card Where the chapter puts it
Two workers did the same work, or part of the task was covered by nobody Duplicated work and gaps First seam
Finished pieces do not fit together Conflicting implicit assumptions Second seam
A caveat or detail vanished on the way up The game of telephone Third seam
A return carries an instruction nobody gave The laundered return Fourth seam
The synthesis is confident and wrong Compounding error across agents Under all four seams
Parallel branches will not merge The merge wall The chapter’s last section
It worked once and will not repeat Non-reproducibility An operational note
The bill is several times what was expected The multiplied bill Stated with the pattern

Real runs often show more than one symptom, so the tool lets you tick several. The chapter explains why they travel together: “Compounding error does not pause at agent boundaries; it crosses them with its confidence intact.” The compounding error calculator shows how fast a chain of steps loses reliability, and the same multiplication applies across agents.

A card can point you to a seam. It cannot find the failing step for you. The chapter’s advice is to “record every brief and every return, because those crossings are where the bugs live”. The agent bug bestiary covers the single-agent failures you will meet once you open the trace.

What makes a good worker brief?

A good worker brief names four things: an objective, an output format, tool and source guidance, and boundaries. Chapter 11 introduces the list with the sentence “A good delegation gives the worker four things.” It also explains why the brief matters more than any other text in the system. A worker starts with no memory of the project, did not hear the user, and cannot see its siblings.

A worker inside its own context window sees only the memo the orchestrator wrote, while the user's request, the planning, and its sibling workers stay outside, unseen.
Figure 11.2 A delegated worker’s entire world: an empty desk, one memo, and siblings it can never see. The worker is exactly as good as the memo that summoned it. Reuse this diagram

The chapter’s example of a brief that fails is one line, “research the semiconductor shortage”, handed to three workers. The result, in its words: “Three workers, one vague sentence, two identical searches and an uncovered flank.” The chapter then makes the point that turns this from bad luck into a fixable defect: “A brief is the cheapest artifact in the whole system to improve, and the most commonly skimped.”

The brief check on this page looks for the four fields in text you paste. It is a heuristic of the tool’s own. It searches for words that usually mark each field and shows the phrase it matched. It will miss a field written in unusual words, and it will sometimes find one that is not really there, which is why every field has a switch to overrule it. It also flags any brief shorter than 12 words. That threshold is the tool’s, not the book’s.

Passing the check is not the goal. The chapter gives a better test: “write every brief as if it were a ticket for a contractor who has never seen your project and cannot ask questions.” Two buttons load the chapter’s own briefs: the illustrative brief it says would pass, and the one-line brief that failed. The chapter marks the passing brief as illustrative and not a template to copy.

Why do well-briefed workers still produce pieces that do not fit?

Well-briefed workers still produce mismatched pieces because each one makes decisions nobody wrote down, and no worker can see the others’ decisions. This is the chapter’s second seam, and it is the one a better brief does not fix.

The chapter borrows its example from an essay that splits a small game in two. One worker builds the background and another builds the character. Both deliver, in two different visual styles. The chapter generalizes: “every action an agent takes encodes silent decisions (a code style, an API choice, a way of handling the empty case), and parallel workers cannot see each other’s silent decisions.” The cost comes late: “The conflicts surface at integration, which is the most expensive possible moment to discover them”.

The remedy is structural. The chapter calls it “the one structural rule the field has genuinely converged on”, which is to keep writes single-threaded. It adds: “Writes are where conflicting assumptions become permanent.” Read-only work does not have this problem, so it can still be done in parallel.

What is the merge wall?

The merge wall is the point where parallel branches have to become one artifact again and cannot be merged mechanically. Chapter 11 takes the name from a practitioner who ran many coding workers at once and found that splitting the work was easy and joining it was not.

A task fans out to four parallel workers, but their results must pass one at a time through a single merge-queue gate in a wall to become one artifact again.
Figure 11.5 Fan-out is cheap and orderly; the reduce is a wall with one narrow gate. Parallel branches must serialize into single-file writes to become one artifact again. Reuse this diagram

The figure’s caption states it: “Fan-out is cheap and orderly; the reduce is a wall with one narrow gate.” The chapter lists three remedies that survived: a merge queue, triage between work that is parallel by nature and work that is serial, and a rhythm of pausing the swarm when a change touches everything. It then ties the first remedy back to the earlier rule: “a merge queue is single-threaded writes, the rule of this chapter’s third section, enforced at the last possible moment, at integration, when it can no longer be politely declined.”

Why should a worker’s summary be treated as untrusted?

A worker’s summary should be treated as untrusted because it can carry an instruction the worker picked up while reading. The chapter calls this the fourth seam and warns that it is easy to miss. A return reads like a colleague’s note, so it gets more trust than a raw tool result would.

The chapter’s wording: “If any of those things carried an injected instruction, the worker’s calm summary is now the delivery vehicle, laundered into a voice your system was built to trust.” Its rule follows: “a worker’s return is tool output in the sense that matters, untrusted until checked, no matter how much its prose sounds like a colleague’s.”

Two other tools on this site go further on this point. Spot the injection is practice at seeing an instruction hidden in content, and the lethal trifecta audit checks whether an agent that reads untrusted text can also reach private data and send it out.

When should work be split across agents at all?

Work should be split across agents only when its branches can run without seeing each other. Chapter 11 starts from the default, “Start with one agent.”, and then asks one question: “does this work decompose into branches that can run without seeing each other?” It gives three conditions in a single line: “No shared mutable state, no dependence on each other’s choices, no ordering that matters.”

The third part of the tool turns those three conditions into questions. The rule that maps answers to a verdict is the tool’s own. Until the answers point one way, the tool shows the chapter’s starting point, one agent. A fourth question appears after the first yes or after three noes, and the rule then has three outcomes.

  • Orchestrator–worker is on the table. All three answers are no. The tool then asks about one of the chapter’s two supporting signs, whether the information genuinely overflows one window. If it does not, the verdict carries a warning.
  • Keep one agent. Any answer is yes. The chapter’s sentence for this case: “if the steps share state or must agree on silent decisions, keep one continuous context, because splitting it fragments exactly the coherence the task depends on.”
  • One writer, parallel advisers. Any answer is yes, and you answer the fourth question by saying that part of the work is read-only. One agent owns the artifact, and the reading fans out.

The tool never says that splitting is required. The chapter’s own phrase is that the pattern is on the table. The chapter is also direct about the price. It quotes one team’s figures, “agents use roughly four times the tokens of a chat interaction, and multi-agent systems roughly fifteen times.”, and concludes: “Parallelism buys wall-clock speed and breadth of coverage; it never buys efficiency.” Those figures are one team’s report and will not match your system. The should-this-be-an-agent tool covers the earlier choice between a script, a model call, a workflow and an agent.

What should I do after reading the cards?

After reading the cards, you should change one thing at the seam the card names and run the task again several times. Several times, because the chapter warns that “a system of five varies combinatorially, so the seam bug you saw this morning may decline to appear this afternoon.” One clean run after a fix does not show that the fix worked. The pass@k calculator puts numbers on that difference.

The Copy as Markdown button gives you the cards you opened, the brief reading and the split verdict in one block for a post-mortem. The guide to agent patterns places the orchestrator–worker pattern among the others. The full argument, including the two essays that disagree about multi-agent systems and how the chapter reconciles them, is in Chapter 11, Multi-Agent Systems (in the full book).

Questions readers ask

What four things must every worker brief contain?
An objective, an output format, tool and source guidance, and boundaries. Chapter 11 takes the list from a team that shipped a multi-agent research system. The objective is the one thing the worker must produce, the format is the shape of the result, the guidance says where to look, and the boundaries say what belongs to a sibling.
Why does the book say to keep writes single-threaded?
Because writes are where two workers' unstated decisions collide and become permanent. Chapter 11 says reading parallelizes well: search, review and analysis add intelligence without touching the shared artifact. Writing is different, so one agent owns the artifact and the others advise. The chapter calls this the one structural rule the field has converged on.
Why should a subagent's return be treated as tool output?
Because the worker spent its run reading things, and any of them could have carried an injected instruction. Chapter 11 says a return arrives composed and in the house style, so it gets trusted more than raw tool output. Its rule is that a return is untrusted until checked, however much it sounds like a colleague.
Does the tool list the fourteen failure modes of the trace study?
No. Chapter 11 reports that a research group found fourteen modes in three clusters, and it names the clusters and not the modes. The tool keeps to what the chapter itself describes: four seams, the compounding underneath them, the merge wall, non-reproducibility and the bill. The study is listed under sources for readers who want the full taxonomy.
How reliable is the brief check?
It is a rough first pass. The tool searches the pasted text for words that usually mark each field, such as a line starting with the word objective or a sentence starting with do not. It cannot judge whether the objective is the right one. Each field has a switch to overrule the tool, and the chapter's contractor test is the real check.

Sources

  1. Mert Cemri et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657)
  2. Anthropic (2025). How we built our multi-agent research system
  3. Walden Yan (Cognition) (2025). Don't Build Multi-Agents
  4. Walden Yan (Cognition) (2026). Multi-Agents: What's Actually Working