Multi agent system overhead is everything a split adds on top of the work itself. Each worker spends tokens re-reading its own context, briefs and returns cross between agents, turns multiply, and a run gains new places to fail. The cost lands on four lines: tokens, wall-clock time, failure surface and debugging effort.
Three published measurements put numbers on it, and each uses a different baseline. One lab reported multi-agent runs at about 15 times the tokens of a chat interaction (Hadfield et al., 2025). The same lab later reported 3 to 10 times a single agent on equivalent tasks (Phillips et al., 2026). A controlled preprint found an orchestrated design taking 3.8 times the turns of one agent for about the same success rate (Kim et al., version 3, 2026).
Whether to split at all is a separate decision, and the single agent vs multi agent decision table owns it. This companion post gives each study with its denominator, each coordination failure with its trace signature and its Chapter 11 remedy, and a budget you can recompute for your own design.
What does multi agent system overhead cover?
Multi agent system overhead covers four costs that a split adds beyond the task: extra tokens, extra wall-clock time, extra places to fail, and extra effort to find which place failed. Each line has its own unit, so no single multiple summarizes them.
- Tokens. Every worker is a full agent loop that re-reads its own prefix and history on each step. On top of that come the brief going in and the result being merged back.
- Wall-clock time. Parallel branches shorten a run. Waiting, a second wave of workers and a serialized merge lengthen it again.
- Failure surface. Each brief and each return is a seam where information can be dropped, changed or trusted too easily.
- Debugging effort. A failed run leaves several transcripts and the crossings between them.
Chapter 11 of AI Agents, Engineered (in the full book) puts the last two lines in one sentence: “With five agents you own a distributed system, and its bugs live in the seams: a dropped handoff, an assumption two workers silently made differently, a summary that shaved off the one detail that mattered.”
The book’s free glossary narrows the search further. Its entry for orchestrator–worker says that “Nearly every coordination failure traces to a violation at one of the two doors—the brief in, the return out”. A subagent knows only its brief, and the lead knows only what came back. The rest of this post counts what crosses those two doors.
Which published multiple answers which question?
Each published multiple for multi agent system overhead answers a different question, because each has a different baseline. Fifteen times is measured against a chat, 3 to 10 times against a single agent in one lab’s unpublished testing, and 3.8 times counts turns under an equal token budget.
| Figure as published | Measured against | Source and version | Population | What it does not show |
|---|---|---|---|---|
| “about 15× more tokens than chats” (and “about 4×” for one agent) | A chat interaction | Hadfield et al., one lab’s engineering write-up, June 2025 | “In our data”: that lab’s own research product; no sample size given | A ratio against a single agent. Dividing 15 by 4 gives about 3.75, a figure the team did not report |
| “3-10x more tokens than single-agent approaches for equivalent tasks” | A single agent | Phillips et al., the same lab’s blog, January 2026 | “In our testing”; no method, workload or count published | A distribution, or a result independent of the first row |
| 27.7 turns against 7.2 (3.8×), at mean success 0.463 against 0.466 | A single agent, with total reasoning tokens matched | Kim et al., arXiv:2512.08296, preprint, version 3, April 2026 | 260 configurations over six benchmarks and three model families; architecture-level means | A token multiple, since the budget was held equal; production workloads |
| “MAS consumes 4–220× more input (prefill) tokens than its SAS counterpart” | A single-agent counterpart | Gao et al., arXiv:2505.18286, preprint, version 1, May 2025 | “across seven datasets”, from a study of code, math and other tasks run on frameworks built around agent conversation, debate and reflection | A lead with briefed workers. The paper’s Table 3 lists nine datasets with ratios from 1.20 to 220.22, so the sentence and the table disagree; I print the sentence |
I found no independent measurement of a token multiple against a single agent on multi-step agent tasks.
What did the failure taxonomy measure, and what did it leave out?
The main failure taxonomy, MAST, measured which failures appear in benchmark runs of seven open-source multi-agent frameworks and how often each label was applied. It left out overhead, production systems, and any estimate of how often a carefully built orchestrator fails.
The paper is “Why Do Multi-Agent LLM Systems Fail?” by Cemri and colleagues, arXiv:2503.13657. I cite version 3, dated 26 October 2025, the latest I could read in full. It names 14 failure modes in three categories (system design issues, inter-agent misalignment, task verification) and applies them to “1642 annotated execution traces” from coding, math and general agent tasks.
What the percentages are. The paper’s Figure 1 caption says the figures “represent the prevalence of each failure mode and category as observed in our analysis of 1642 MAS execution traces”. The fourteen figures add up to 100.05 by my sum. I therefore read each one as a share of all labeled failure instances. The paper states the denominator no more exactly than that.
Who applied the labels. The taxonomy itself was built by six human experts from “150 traces from five MAS frameworks”. Agreement between humans was then measured on a much smaller set: “three annotators independently and iteratively labeled a total of 15 traces until achieving high inter-annotator agreement (κ = 0.88)”. The method section describes three rounds of five traces, with 0.88 the average “in the final rounds”.
For the scale-up the authors used a model judge, which the paper reports at “accuracy 94%, Cohen’s Kappa of 0.77” against human labels. That check ran “on a held-out set from our IAA studies”, whose size the paper does not state; the human-annotated set it released from those studies holds 21 traces.
By my count from the paper’s Table 1, 210 traces carry human annotations, 30 per framework. The other 1,432, or 87%, carry judge labels only.
Which systems. Two frameworks that simulate the roles of a software company (MetaGPT and ChatDev, named here as dated 2025 examples) supply 760 of the traces, 46% by my count. A third, which the paper describes as a general framework for building agents, supplies another 597, or 36%. A reader who runs a lead with briefed workers is looking at a different architecture from most of the data.
What the authors disclaim. They write “we do not claim MAST is exhaustive”, and they say the per-system performance figures were “measured on different benchmarks, therefore they are not directly comparable”. The paper’s text gives failure rates from 41% to 86.7% across seven systems, while the caption of the figure that plots them says six systems. I would quote that range only with the mismatch attached.
Overhead. Version 2 of the paper stated the gap plainly: “a significant prevalence of inefficiencies in MAS traces, which MAST currently does not include by design”. Version 3 dropped that section. A run that wastes turns and still reaches the right answer has no label in this taxonomy.
Which version do the popular percentages come from?
The category shares most often repeated come from version 2, dated 22 April 2025, which describes “over 200 tasks”. That version prints “FC1: 41.77%, FC2: 36.94%, FC3: 21.30%”, and the middle figure is the source of the “37% of failures are inter-agent misalignment” line.
Version 3 gives no category shares in its body text. Its per-mode figures sum to 44.2, 32.35 and 23.5 for the three categories (my sums), on a dataset several times larger. Individual modes moved as well: “fail to ask for clarification” was 11.65% in version 2 and is 6.80% in version 3.
Which coordination failures should you look for in a trace?
Look for thirteen, each with a signature you can search for in a trace and a remedy from Chapter 11. In the table, “MAST v3” means a share of labeled failures in the 1,642 benchmark traces described above, 87% of them labeled by the judge alone. Where the evidence holds no prevalence figure, the cell says so.
The mapping from the paper’s modes to the book’s failures is mine. The last column says where the fix lives, and the buttons above the table filter on it.
For loops and ping-pong, the book’s remedies are structural. Its routing rule is “if you can write the subtask list before seeing the input, use the workflow”: a handoff whose target follows from a rule is a branch in code. Its termination rule is to “cap the loop, because generator and critic can circle each other indefinitely”.
| Failure | How it shows in a trace | Measurement, with its denominator | Chapter 11 remedy | Where the fix lives |
|---|---|---|---|---|
| Duplicated work and gaps | Sibling subtrees with near-identical tool calls; a sub-question with no span at all | No prevalence figure found. In one lab’s 2026 experiment with peer agents, 18 of 30 same-model agents chose the same branch name | One objective per worker, plus boundaries that name what belongs to a sibling | brief |
| A guess where a question was needed | A worker proceeds on a value that appears in no brief and no tool result | MAST v3: fail to ask for clarification 6.80% | Write the brief for a contractor who has never seen the project and cannot ask | brief |
| Loops and repeated steps | The same delegation or tool span repeating; A to B to A handoffs; a turn count far above healthy runs | MAST v3: step repetition 15.7%, unaware of termination conditions 12.4% | “cap the loop”; a fixed output format, so that “done” has one shape; routing by rule moved into code | structure |
| Too many workers for the job | A fan-out wider than the number of independent subtasks | No prevalence figure found. One lab says its early system was “spawning 50 subagents for simple queries” (2025, one system) | Explicit scaling rules in the orchestrator’s prompt | structure |
| Conflicting implicit assumptions | Two artifacts that each pass their own check; the first failure appears at integration | No measurement found | “keep writes single-threaded”: one agent owns the artifact | structure |
| The merge wall | Rebases that fail; work redone on a baseline that moved | No measurement found; Chapter 11 relies on one practitioner’s account | A merge queue, triage of serial work, a pause to serialize | structure |
| Operator overload | Finished results waiting unreviewed while new agents start | No measurement found; Chapter 11 cites practitioners’ self-reports | A work-in-progress limit and a visible status board | structure |
| Detail lost at the return | The worker’s transcript holds a caveat that its return, or the lead’s restatement, lacks | MAST v3: loss of conversation history 2.80%, ignored other agent’s input 1.90%, information withholding 0.85% | Persist large artifacts and pass a reference; start with return-values-only | return contract |
| The trusted return | Text that first appears in a worker’s tool result comes back as an instruction the lead acts on | No measurement found | Treat a return as tool output, “untrusted until checked” | verification |
| Derailment and silent drift | The final answer addresses a different subject from the first user message while every span reports success | MAST v3: disobey task specification 11.8%, task derailment 7.40% | Resource the synthesis step; compare the result with the original request | verification |
| Unchecked or wrongly checked merge | Synthesis follows the returns with no check span between, or a check that only compiles | MAST v3: incorrect verification 9.10%, no or incomplete verification 8.20%, premature termination 6.20% | A read-only reviewer with a fresh window, the finished work and the standard to judge it by | verification |
| An error that crosses a boundary | A worker’s wrong claim reappears word for word in the lead’s answer | Kim et al. v3: error amplification 17.2 for independent workers against 4.4 with an orchestrator; after controls it does not predict performance (p = 0.658) | A check at every return; the compounding arithmetic of Chapter 2 | verification |
| A seam you cannot find | No record of what was sent or returned; a failure that does not repeat | Zhang et al., 2025: best automated method names the agent 53.5% of the time and the step 14.2%, on logs from 127 systems | “record every brief and every return” | logging |
Three of the paper’s modes are missing from the table because I could not tie them to one seam. They are reasoning-action mismatch (13.2%, the second largest), conversation reset (2.20%) and disobeying a role specification (1.5%). The post on AI agent failure modes and their trace signatures covers the shapes one agent produces alone, and the agent bug bestiary asks whether a wrong turn arrived from another agent.
How should you read the shares in that table?
Read each MAST share as “of the failures labeled in these benchmark runs, this fraction received this label”. A share tells you where to look first in a failed trace. It carries no information about how often your own system will fail, or about what the failure will cost you.
Two findings from the same paper limit what a remedy can promise. The authors write that “Solutions focused on context or communication protocols are often insufficient” for the inter-agent category.
They also tested two interventions. On one role-simulating framework with 32 tasks, success went from 25.0% to 34.4% with an improved prompt and to 40.6% with a new topology: 8, then 11, then 13 tasks by my arithmetic. That is the largest gain of the four benchmark columns in the paper’s Table 5; the other three moved by at most five points.
Verification deserves the same caution. The paper reports that “many existing verifiers perform only superficial checks, despite being prompted to perform thorough verification”. A reviewer agent that reads a summary of the work has the same blind spot as the lead that wrote the summary.
What did the controlled comparison find?
The controlled comparison found that coordination multiplied turns without raising average success. Under a matched budget, a centralized design averaged 27.7 turns and a success rate of 0.463, against 7.2 turns and 0.466 for a single agent. The average hides a wide spread by task.
That study is Kim and colleagues, “Towards a Science of Scaling Agent Systems”, a preprint whose version 3 is dated April 2026. It ran 260 configurations across six benchmarks, five architectures and three model families, with “All systems matched for total reasoning tokens”. More agents therefore meant less reasoning per agent, a different condition from a team that is simply given more budget.
By task, the relative change ran “from +80.8% on decomposable financial reasoning to -70.0% on sequential planning”. The authors also report that tasks “where single-agent performance already exceeds 45% accuracy experience negative returns from additional agents”. Both findings rest on six benchmarks, two of them run on 20-instance subsets, so I treat the 45% as a prompt to measure a baseline first.
A checking step did measurable work. Architectures with a verification step, either an orchestrator that cross-checks sub-agent outputs or rounds of peer debate, “achieve 22.7% average error reduction (95% CI: [20.1%, 25.3%])” in factual error rate. Independent workers showed none. That is the published support for the verification rows above, within one study’s benchmarks.
Why are seam bugs slow to find?
Seam bugs are slow to find because the evidence is spread over several transcripts, and automated attribution is still weak. In the “Who&When” benchmark of Zhang and colleagues (arXiv:2505.00212, version 3, June 2025), “The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps”.
The dataset holds “failure logs from 127 LLM multi-agent systems”, mostly agent teams generated by an algorithm, and the paper says those generated teams all run on one 2024 model version. Its three human annotators spent 30.9, 30.2 and 23.2 hours on the labels and were uncertain about 15% to 30% of them.
Chapter 11 adds a second reason, which no study I found has measured: “a system of five varies combinatorially, so the seam bug you saw this morning may decline to appear this afternoon”. A stored copy of each brief and each return is the only part of the run that stays put.
What does a split cost? An illustrative budget
In this illustration of multi agent system overhead, a lead with three workers spends 226,700 tokens, against 80,000 for a ten-step single agent: 2.83 times. Every number in this section is my own arithmetic on round inputs. None of it is a measurement, and the taxonomy above cannot supply one, because it counts failures and ignores wasted work.
The inputs are the napkin numbers of Chapter 19 (in the full book), which the agent cost-per-task estimator uses as defaults. The lead’s fixed prefix is 3,200 tokens. On each step an agent writes 300 tokens and reads a 700-token tool result, so its history grows by 1,000 tokens a step.
A worker’s prefix is 2,000 tokens, its brief 1,500 tokens in, and merging its result 1,000 tokens out. There are no retries and no caching.
An agent re-reads its prefix and its whole history on every step. Over n steps its input is prefix × n + 1,000 × n(n − 1)/2, and its output is 300 × n.
- One agent, 10 steps. Input: 3,200 × 10 + 1,000 × 45 = 77,000. Output: 300 × 10 = 3,000. Total 80,000.
- One worker, 8 steps. Input: 2,000 × 8 + 1,000 × 28 = 44,000, plus the 1,500-token brief = 45,500. Output: 300 × 8 + 1,000 for the merge = 3,400.
- Lead plus three workers. Input: 77,000 + 3 × 45,500 = 213,500. Output: 3,000 + 3 × 3,400 = 13,200. Total 226,700, which is 2.83 times the ten-step agent.
- One agent doing all 34 steps in one context. Input: 3,200 × 34 + 1,000 × 561 = 669,800. Output: 300 × 34 = 10,200. Total 680,000, which is 3.0 times the split.
| Line | One agent, 10 steps | Lead (10 steps) with 3 workers (8 steps each) | One agent, 34 steps, one uncompacted context |
|---|---|---|---|
| Input tokens | 77,000 | 213,500 | 669,800 |
| Output tokens | 3,000 | 13,200 | 10,200 |
| Total tokens | 80,000 | 226,700 | 680,000 |
| Wall-clock at an assumed 6 s a step | 60 s | 108 s | 204 s |
| Model-driven steps | 10 | 34 | 34 |
| Contexts to read after a failure | 1 | 4 | 1 |
| Seams (briefs in, returns out) | 0 | 6 | 0 |
| Chance every step and seam is right (assumed 98% a step, 95% a seam, independent) | 82% | 37% | 50% |
The estimator below opens on the lead with three workers and reproduces its tokens: 213,500 in, 13,200 out, 226,700 in total. Set the workers to zero for the ten-step agent, then raise the steps to 34 for the long run. The retry overhead is set to zero here; the tool’s usual default is 10%. The estimator computes the three token rows only, and the other five rows stay hand arithmetic.
With JavaScript on, the Agent cost-per-task estimator runs here, filled in with the example from this post.
Runs in your browser; nothing is sent anywhere. Open the Agent cost-per-task estimator on its own page to share a result by link.
What happens to wall-clock time?
Wall-clock time falls only against the long single run. At an assumed 6 seconds a step, the ten-step agent takes 60 seconds and the 34-step agent 204 seconds. The split takes the lead’s 60 seconds plus the slowest worker’s 48, or 108 seconds, if all three workers run at once and the lead waits one time.
Run the workers one after another and the split takes 60 + 3 × 48 = 204 seconds, the same as the long run. One lab notes that multi-agent systems “often take longer overall than single-agent systems because of the sheer increase in total computation” (Phillips et al., 2026).
What is the chance that every participant succeeds?
Under independence, the chance that everything goes right is the product of every step’s and every seam’s success rate. Assume 98% a step and 95% a seam, both placeholders.
Ten steps give 0.9810 ≈ 82%. Thirty-four steps give 0.9834 ≈ 50%. Thirty-four steps and six seams give 0.50 × 0.956 ≈ 0.50 × 0.74 ≈ 37%.
Treat that row as exposure. It assumes that errors are independent, that any error is fatal and that nothing checks a return, and real runs break all three assumptions in both directions. A lead that verifies returns recovers some of the loss. The post on why agent errors compound explains why chains often do worse than the product, and Chapter 2 (free to read) has the base arithmetic.
The compounding error calculator reproduces the 50% cell from a 98% step rate over 34 steps. Chapter 11’s sentence about compounding errors in AI agents is the reason seams get their own factor: “Compounding error does not pause at agent boundaries; it crosses them with its confidence intact.”
Does a multi-agent split always cost more tokens?
No. In the two reports that publish a token multiple against a single agent (Phillips et al., 2026; Gao et al., 2025), the multi-agent version spent more. One vendor benchmark measured the reverse, and the budget above shows one condition under which the comparison flips: a single agent that carries all the work in one context that is never compacted.
Chapter 11 states the general claim without hedging: “More agents genuinely means more tokens, full stop, and on that measure parallelism only ever loses.” That holds for the budget’s comparison with the ten-step agent, where the workers’ 24 steps are coverage added on top of the lead’s ten. It also holds for the measured multiples, which compare systems as their builders ran them. The chapter’s next point belongs beside it: workers can run on a cheaper model tier, so “the token count still multiplies while the bill does not have to”.
The 34-step line makes a different assumption: the lone agent must take the same 34 steps, and it re-reads all of its history on each one. History re-reading grows with the square of the step count, so 34 steps in one window cost far more than four short windows. With these inputs the single agent’s total passes the split’s 226,700 at step 19, where it reaches 237,500.
Real single agents trim and compact their context, and they often need fewer steps than a team, since nothing has to be briefed or merged. So 680,000 is an upper bound, and the reversal is a property of long, uncompacted runs. A forum commenter made the same point in one line: “Subagents are expensive but they scale way closer to O(n) than O(n^2)” (jmalicki, Hacker News, September 2026).
One vendor’s benchmark of its own library reports a related reversal, through the prefix. On a modified customer-support benchmark with unrelated domains added, “the single agent uses consistently more tokens as the number of distractor domains grows, while supervisor and swarm remain flat” (Fu-Hinthorn, 2025). The page prints no values. Read off its chart, the single agent spends about 100,000 tokens with one distractor domain and about 220,000 with six, while both multi-agent designs stay between roughly 70,000 and 85,000 at every count from one to eight. With no distractor domain the chart plots only the single agent, at about 68,000, so it does not show what a split costs when one agent’s context is small.
The same post names the price for its supervisor design: “Most of the performance issues for the supervisor architecture came from the ‘translation’” of what a sub-agent had returned. It adds that its other design, which it calls a swarm, avoids this, because there a sub-agent answers the user directly.
So the subagents vs single agent question on tokens has two honest answers. Against the agent you run today, a split adds tokens. Against the much longer run it replaces, multi agent system overhead can turn negative on the token line. You only know which case you are in after you compute both.
When is one agent the cheaper design, and when does the split win?
One agent is cheaper whenever the workers’ steps are added work, the task is already within its reach, or the output is one shared artifact. The split wins when independent, read-heavy branches would otherwise sit in one long context, or when wall-clock time is the binding constraint.
| Situation | Cheaper design | The assumption behind it |
|---|---|---|
| The task fits in a short run and workers would add coverage you do not need | One agent | The workers’ steps are extra work, as in the budget’s comparison with the ten-step agent |
| Dozens of steps of independent reading would pile up in one window | The split | The single agent would take the same steps and would not compact; measure a compacted baseline before you believe it |
| One agent already succeeds on most of the tasks | One agent | The saturation finding of Kim et al. (v3) transfers from its six benchmarks to your work |
| The output is one shared artifact, such as a codebase | One writer, with readers fanned out if needed | Conflicts found at merge cost more than the serial time saved |
| Latency binds and the branches are independent | The split | One wave, full parallelism, and no serialized merge at the end |
| The split follows job titles: planner, implementer, tester, reviewer | One agent | One lab’s single experiment generalizes: there, “the subagents spent more tokens on coordination than on actual work” |
The last row comes from one experiment that the lab reports without numbers (Phillips et al., 2026), so weigh it as a warning. The latency row has its own counterexample: in one vendor’s flat design with a shared file and locks, “Twenty agents would slow down to the effective throughput of two or three, with most time spent waiting” (Lin, 2026).
For the shared-artifact row, Chapter 11 has a name for the place where parallel writers collide, the merge wall, and a drawing of the mechanism.
The book’s comment on the merge queue applies to the whole row: “The wall is the same law you met at delegation time, now collecting with interest.” Serialized writes are cheaper when you choose them at design time.
Chapter 19 (in the full book) supplies the test for the two rows where the split wins: “the multiplied bill is worth paying exactly when the task’s value clears it and breadth or speed is what wins”. Apply the same test one level up and you learn when not to use AI agents at all: when the task’s value does not clear even one agent’s bill.
What should you check before committing to a split?
Check ten things, and treat the second and third as gates. If the branches need to see each other, or if two agents would write to one artifact, restructure before you tune anything else. The remaining eight items are work to finish before the first production run.
- Baseline. One agent has been measured on the same tasks: success rate, tokens per task, wall-clock time.
- Gate: independence. The branches that run at the same time can run without seeing each other: no shared mutable state, no dependence on each other’s choices. A reviewer or check that runs after a branch is a sequential step, and this gate does not apply to it.
- Gate: one writer. Each artifact has one writer, and a codebase with shared modules counts as one artifact. If two agents must write to it, their writes pass through a merge queue you have budgeted for.
- Briefs. Every brief names the objective, the output format, the tools and sources, and the boundaries.
- Returns. Every return has a fixed shape, and large artifacts are stored and passed by reference.
- Logging. Every brief and every return is recorded, keyed by run and worker, inside one trace tree.
- Checks. A check sits between each return and synthesis, and it reads the artifact itself.
- Budgets. Step, token and time limits cap every worker and the fan-out, and any loop between two agents has a hard cap in code.
- Token budget, twice. The split is priced against the short single run and against the long uncompacted one.
- Exit. You have written down the result that would make you remove the agents.
The independence gate is the book’s own question from Chapter 11: “does this work decompose into branches that can run without seeing each other?” The list and the budget assume a lead whose workers report back. In a handoff design, where a specialist answers the user directly, nothing returns, so apply the Returns and Checks items to the handoff message and to the specialist’s final answer. The order of work for how to build a multi-agent system from nothing is a separate subject, and so is the wider set of agent patterns a split belongs to.
Do the table and the checklist agree on two real designs?
They agree on both designs I tried. I ran each one through the failure table and then through the checklist, looking for a case where the two would give different advice.
Three parallel research workers with independent sub-questions. Both gates pass: the sub-questions do not depend on each other, and only the lead writes the report.
The table points to three rows for this topology: duplicated work (brief), detail lost at the return (return contract) and an unchecked merge (verification). The checklist’s brief, return and check items cover the same three. The budget is the worked one above: 2.83 times the tokens of a short single run, and 108 seconds against 204 if the alternative is one long run.
Two workers editing the same codebase. Both gates fail. The workers depend on each other’s silent choices, and one artifact has two writers.
The table’s matching rows are conflicting implicit assumptions and the merge wall, and both remedies are serialized writes: one owner chosen up front, or a merge queue at the end. The checklist’s one-writer gate offers the same two options. I would take the first: one writing agent, with the second agent kept read-only as a reviewer.
I did find one wrinkle. Two workers editing disjoint paths that share no module pass the one-writer gate as two artifacts, and then the independence gate decides. If either worker’s code calls the other’s, that gate fails, and the answer is again one writer.
Where does this analysis stop?
It stops at the edge of the published evidence on multi agent system overhead, which is thin in three places. The token multiples against a single agent come from one lab. The taxonomy describes benchmark runs of 2025 frameworks. Four of the thirteen failure rows have no measurement at all.
The budget is arithmetic on placeholders. Six seconds a step, 98% a step and 95% a seam are round numbers chosen for the example, and the formula ignores compaction, caching, retries and a second wave of workers.
The studies will also age: the taxonomy’s authors write that “MAS failure is not merely a function of challenges in the underlying model”. Gao and colleagues report that “the benefits of MAS over SAS diminish as LLM capabilities improve”. Neither claim says that coordination becomes free. Expect the shares to change and the two doors to stay.
I left out per-handoff latency because I found no figure with a checkable source.
The takeaway
Before you orchestrate multiple AI agents, compute the budget on three lines and open the trace at the two doors. The overhead is measurable in your own system even where the published numbers are thin. Count the steps, the contexts, the seams and the writers, then compare with one agent on the same tasks. Chapter 11 asks that “the bill gets measured rather than assumed”, and the same holds for the failure surface.
Each coordination failure in the table comes with one more thing you can check: a duplicate span, a caveat missing from a return, a merge with no check before it.
Chapter 11, “Multi-Agent Systems,” is in the full book, with the coordination failure field guide, the subagent archetypes and the merge wall in depth. The glossary and Chapters 1 and 2 are free to read online, and you can see the formats.
Questions readers ask
- Is it better to put several agents' functions into a single agent to avoid multi-agent overhead?
- Often, yes. A controlled preprint (Kim et al., arXiv:2512.08296, version 3, 2026) found that an orchestrated design reached about the same average success as one agent (0.463 against 0.466) while taking 3.8 times the turns, across six benchmarks under a matched token budget. The exception is work that splits into independent, read-heavy branches, where the same study measured gains of up to 80.8% on one financial reasoning benchmark.
- How many more tokens does a multi-agent system use than a single agent?
- One lab reported in January 2026 that its multi-agent implementations typically used 3 to 10 times the tokens of a single agent on equivalent tasks; it published no method or sample. The better-known 15 times figure, from the same lab in June 2025, is measured against a chat interaction. I found no independent measurement of a token multiple against a single agent on multi-step agent tasks.
- Do subagents scale linearly or quadratically in token cost?
- Both terms are present. Each agent's input grows with the square of its own step count, because it re-reads its history on every step, while adding workers of a fixed length adds cost roughly linearly. In this post's illustration, one agent carrying 34 steps in one uncompacted context reads 669,800 input tokens, and a lead with three short workers doing the same 34 steps reads 213,500.
- Why do agents loop or hand a task back and forth, and how do you stop it?
- The largest single share in the MAST taxonomy (version 3) is step repetition, at 15.7% of labeled failures in benchmark runs of seven open-source frameworks, and unawareness of termination conditions adds 12.4%. Chapter 11 of AI Agents, Engineered tells you to cap a loop between a generator and its critic, to size the fan-out with written rules, and to use a fixed workflow wherever the subtask list can be written in advance. A step, token and time budget per worker, enforced in code, is the stop that a prompt cannot argue with.
- How do you find which agent caused a multi-agent failure?
- Log every brief and every return, keyed by run and worker, inside one trace tree. Automated attribution is weak: in a 2025 benchmark of failure logs from 127 multi-agent systems (Zhang et al., arXiv:2505.00212), the best method named the responsible agent 53.5% of the time and the failing step 14.2% of the time, and three human annotators each spent between 23 and 31 hours labeling.
Sources
- Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657; version 3 of 26 October 2025 is the one cited, version 2 of 22 April 2025 for the comparison)
- Yubin Kim, Ken Gu, Chanwoo Park, et al. (2025). Towards a Science of Scaling Agent Systems (arXiv:2512.08296; preprint, version 3 of 8 April 2026)
- Shaokun Zhang, Ming Yin, Jieyu Zhang, et al. (2025). Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems (arXiv:2505.00212; version 3 of 2 June 2025)
- Mingyan Gao, Yanzi Li, Banruo Liu, et al. (2025). Single-agent or Multi-agent Systems? Why Not Both? (arXiv:2505.18286; preprint, version 1 of 23 May 2025)
- Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford (Anthropic) (2025). How we built our multi-agent research system (engineering blog, 13 June 2025)
- Cara Phillips, with contributions from Paul Chen, Andy Schumeister, Brad Abrams and Theo Chu (2026). Building multi-agent systems: when and how to use them (vendor blog, 23 January 2026)
- Anthropic (2026). Patterns and problems in emerging multi-agent systems (research post, 13 August 2026)
- Will Fu-Hinthorn (LangChain) (2025). Benchmarking Multi-Agent Architectures (vendor blog, 10 June 2025)
- Wilson Lin (Cursor) (2026). Scaling long-running autonomous coding (vendor blog, 14 January 2026)
- Hacker News commenter storus (2026). Forum comment asking whether one agent should absorb several agents' functions (2 February 2026)
- Hacker News commenter jmalicki (2026). Forum comment on subagent cost scaling (30 September 2026)
- GitHub user lan3344, microsoft/autogen issue tracker (2026). Issue #7487: goal drift that every agent reports as success (29 March 2026)