The single agent vs multi agent decision has a short answer: start with one agent. Split the work across several only when it decomposes into branches that can run without seeing each other, and whose results can be merged more cheaply than they were produced. Everything else, from the token bill to the failure modes, follows from that one test.
I hold that position because the evidence on both sides now exists in writing. One team reported that its multi-agent research system used about 15 times the tokens of a chat interaction (Hadfield et al., 2025). A controlled study of 260 agent configurations found multi-agent designs ranging from 80.8% better than a single agent on decomposable financial reasoning to 70% worse on sequential planning (Kim et al., 2025). Same idea, opposite outcomes.
The variable is the shape of the work, the same lesson as the agents vs workflows decision rule one level up, and this post gives you a single agent vs multi agent decision table for reading it, plus a worked example you can reuse in a design review.
Single agent vs multi agent: what is the actual difference?
A single agent runs one loop in one continuous context: every step sees everything before it. A multi-agent system splits the work across several loops with separate contexts, usually an orchestrator that decomposes the task and workers that each handle a piece and return a distilled result. The split trades shared context for parallelism.
That trade is the whole debate. Two essays published within days of each other in June 2025 framed it. One, from a team building a coding agent, argued that “running multiple agents in collaboration only results in fragile systems,” because “the decision-making ends up being too dispersed and context isn’t able to be shared thoroughly enough between the agents” (Yan, 2025). The other, from a team building a research product, reported that its multi-agent design beat a single agent by 90.2% on its own internal research evaluation (Hadfield et al., 2025).
Chapter 11 of AI Agents, Engineered (in the full book) reads both essays against each other and concludes that each team was reporting accurately from its own domain. Research is read-heavy and composes by juxtaposition; coding mutates one shared artifact and composes by merge. The chapter’s verdict on multi-agent design is one line I keep coming back to: “It is an architecture with a habitat.”
What does a multi-agent system look like in practice?
In practice, almost every production multi-agent system is the orchestrator–worker pattern: a lead agent receives the task, decides at runtime how to split it, hands each piece to a worker running its own agent loop in a fresh context, then synthesizes the returns into one answer. Peer-to-peer swarms negotiating among themselves are rare outside demos.
The word that matters is runtime. If you can list the subtasks before the input arrives (run these three checks on every ticket), you don’t need an orchestrator at all; a fixed fan-out in ordinary code does the job.
The book puts the rule plainly: “if you can write the subtask list before seeing the input, use the workflow. It is simpler, cheaper, and easier to debug. Reach for an orchestrator only when producing the list itself takes intelligence.” The same distinction appears in the essay that popularized the pattern, which describes a central model that “dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results” (Schluntz and Zhang, 2024).
Each worker is a subagent: a complete loop with its own context window, driven by the same kind of agent harness that runs a single agent, which means it can spend a great many tokens on dead ends and hand back a short summary while the mess stays sealed inside. That is the real prize, more than speed. Each fresh desk postpones context rot, the slow degradation of a model’s attention as its window fills. The book files this under isolate, the fourth verb of write, select, compress, isolate, and observes that “A subagent is a context decision before it is an organizational one.” For the craft of designing the pattern itself, see the companion post on writing the orchestrator–worker brief, and the glossary entry for orchestrator–worker.
Why does the worker’s brief decide whether it works?
The brief decides it because a worker knows nothing except what the brief says. It did not hear the user, did not see the plan, and cannot see its siblings. A vague brief produces overlapping work and uncovered gaps, and the failure is authored upstream, in text the orchestrator wrote, before any worker runs.
The research team quoted above reports exactly this. Its early orchestrator handed workers short instructions such as “research the semiconductor shortage,” and in one case a worker explored the 2021 automotive chip crisis while two others duplicated work on current supply chains, “without an effective division of labor”. Their fix was a four-part brief: an objective, an output format, guidance on tools and sources, and explicit boundaries telling each worker what belongs to a sibling. The same team admits its early system was capable of “spawning 50 subagents for simple queries” until explicit scaling rules went into the orchestrator’s prompt.
My own test in reviews is the one the book recommends: “write every brief as if it were a ticket for a contractor who has never seen your project and cannot ask questions.” If a competent stranger could execute it from the text alone, a worker probably can. The chapter closes the section with the sentence that belongs on every orchestrator’s prompt: “The worker is exactly as good as the memo that summons it.”
What does a multi agent system cost compared with one agent?
A multi-agent system costs more tokens than a single agent, always, because every worker re-reads its own instructions, tool results and history on every step. One team’s point-in-time measurement put agents at about 4 times the tokens of a chat interaction and multi-agent systems at about 15 times (Hadfield et al., 2025).
Most pages that quote the 15× figure miss what it is measured against. It compares a multi-agent run to a chat, not to a single agent.
Divide one multiplier by the other (15 / 4 ≈ 3.75) and the same team’s numbers suggest a multi-agent run costs roughly three to four times what a single agent would on comparable work. That derivation inherits every limit of the original (one workload, one moment, one team’s systems), so treat it as an order of magnitude to check against your own traces, not a constant. For the single agent vs multi agent budget question, it is still a better anchor than the bare 15×.
The book’s summary is the sentence to bring to the budget meeting: “Parallelism buys wall-clock speed and breadth of coverage; it never buys efficiency.” Two things soften the bill without changing that truth. First, tokens are not a fixed price, and nothing requires workers to run on the same model as the orchestrator; spend the expensive tier on decomposition and synthesis and let a cheaper tier do the bulk reading. Second, the same team found that token usage alone explained 80% of the performance variance on one browsing benchmark, which means some of what multi-agent “buys” is simply more total reasoning applied to the problem. To price a design before building it, the agent cost-per-task estimator lets you plug in your own step counts and multipliers.
| Cost line | Single agent | Multi-agent (orchestrator–worker) |
|---|---|---|
| Tokens | One context, re-read each step | Every worker re-reads its own context; about 3–4× a single agent by my derivation from one team’s 4× and 15× (both against chat) |
| Wall-clock time | Sequential | Parallel branches run at once; can be far faster on broad work |
| Context quality | Degrades as one window fills | Each worker starts clean; the lead sees only summaries |
| Debugging | One linear trace | Bugs live in the seams between briefs and returns |
| Reproducibility | Varies run to run | Varies combinatorially across workers |
| Operator attention | One stream to supervise | Several streams, plus the merge |
How do multi-agent systems fail?
Multi-agent systems fail mainly at the seams between agents rather than inside any one of them: duplicated work and gaps from vague briefs, conflicting silent assumptions between parallel workers, detail lost as summaries pass upward, and errors that cross agent boundaries with their confidence intact. These failures now have a published taxonomy.
The best-known map is MAST, the Multi-Agent System Failure Taxonomy of Cemri and colleagues (2025). They built it from close analysis of 150 traces with expert annotators (inter-annotator agreement κ = 0.88), then applied it to a dataset of more than 1,600 annotated traces across seven popular multi-agent frameworks. It identifies 14 failure modes in three clusters: system design issues, inter-agent misalignment, and task verification. The paper opens with the observation that multi-agent systems’ “performance gains on popular benchmarks are often minimal.” The full catalog, mode by mode, is the subject of the companion post on multi-agent system overhead and its coordination failure modes.
The deepest of these is the one the “don’t build” essay named: “Actions carry implicit decisions, and conflicting decisions carry bad results.” Two workers asked to build a game background and a game character can both succeed and still deliver pieces in different visual styles that do not compose. No brief enumerates every silent decision in advance, so the conflict surfaces at integration, the most expensive moment to find it.
Errors also compound across agents the way they compound across steps. Chapter 2 (free to read) runs the arithmetic: at 95% per-step reliability, a twenty-step chain succeeds about 36% of the time. A worker that misread its brief returns a wrong answer in the same fluent register as a right one, and synthesis weaves it in. The post on why agent errors compound works through the math, and the compounding error calculator lets you run your own numbers.
Architecture changes how badly. In the controlled study by Kim and colleagues, independent workers with no central check showed a trace-level error amplification of 17.2, against 4.4 for systems where an orchestrator reviewed outputs before aggregating them. The authors caution that these are measures of the extra computational work associated with coordination failures, not the odds of a wrong final answer (MIT Media Lab overview). The direction still matters: a verification step at the handoff is worth paying for.
What is the merge wall?
The merge wall is the point where parallel workers, each starting from the same baseline, try to combine their changes and find that the baseline has moved so far that their work no longer applies. Steve Yegge named it in late 2025, after running a swarm of eight coding workers through an orchestrator he built himself (Yegge, 2025).
He frames agent swarming as a MapReduce-style job. The map phase fans out beautifully. The reduce phase is where the analogy breaks, because classic reducers merge mechanically (sum the counts), while “with agent swarming, the reduce phase is a nightmare; it’s the exact opposite, in fact: it can be arbitrarily complicated to merge the work of two agents.” His example is the limit case: one worker deletes a subsystem while another is busy changing it.
His remedies are worth copying. Use a merge queue: “You need to serialize the rebases, and give each worker enough context, and context-window space, to fully merge their work into the new baseline.” Triage the work, because “Some work is inherently parallel, and some work is inherently serial.” And expect a rhythm in which a project is swarmable for a while, then needs every worker paused during a sweeping change such as a directory restructure.
A merge queue is the field’s one point of convergence in disguise. The coding team that wrote “don’t build multi-agents” published a follow-up ten months later describing what was actually working: “setups where multiple agents contribute intelligence to a task while writes stay single-threaded” (Yan, 2026). Reads parallelize; writes serialize. One agent owns each artifact, and everyone else advises.
When does multi-agent pay? The decision table
Multi-agent pays when the subtasks can only be found at runtime, the branches are independent, the work is mostly reading, the information overflows one window, each return can be checked cheaply, and the results combine by assembly. If any row lands in the right-hand column, start with a single agent and revisit with measurements.
| Question | Multi-agent tends to pay when… | A single agent tends to win when… |
|---|---|---|
| Can you list the subtasks before seeing the input? | No: finding the list takes judgment | Yes: use a fixed fan-out in code, or one agent |
| Can the branches run without seeing each other? | Yes: no shared state, no ordering | No: steps share state or depend on each other’s choices |
| Is the work mostly reading or writing? | Reading: research, review, evaluation sweeps | Writing to one shared artifact, such as most code changes |
| How do results combine? | Assembly: rows in a table, sections in a report | Integration: a merge where seams are defects |
| Does the information overflow one window? | Yes: each worker can digest a slice | No: one agent can hold it all and stay coherent |
| Can you check each worker’s return cheaply? | Yes: a schema, a citation check, a test | No: verifying a return costs as much as producing it |
| How good is a strong single agent already? | Weak on this task, with room to improve | Already strong; one study saw gains vanish above roughly 45% single-agent success |
| Is wall-clock time worth multiplied tokens? | Yes: latency is the binding constraint | No: budget binds harder than time |
| Who orchestrates, and can they absorb the output? | A model, or a person with spare review capacity | You, already at your review limit |
The capability row comes from the Kim study, whose authors describe a saturation threshold “near 45% single-agent success” and are careful to call it “an empirical selection rule, not a universal cutoff” (MIT Media Lab overview). I read it as a reminder to measure the single-agent baseline first rather than as a number to design around. If the decision is one step earlier (whether the task needs an agent at all), the Should this be an agent? decision tool covers that question.
Worked example: one team, two proposals
To see the table at work, take one platform team bringing two multi-agent proposals to the same design review. Run through the rows, the first proposal earns its multi-agent design and the second does not, even though both sound equally “parallel” on a slide. All numbers below are illustrative.
Does the contract review split cleanly?
Proposal A: before a renewal cycle, review 40 supplier contracts and flag any clause that allows a unilateral price change, with the clause text and its location. The subtasks are not fully known in advance (some contracts reference appendices and master agreements that must be found first), so an orchestrator earns its place.
Each contract can be read without seeing any other. The work is pure reading. The output is a table with one row per contract, which combines by assembly. And each row can be spot-checked cheaply against the cited clause.
The information also overflows one window. Say each contract and its appendices come to about 15,000 tokens; 40 of them is 600,000 tokens; even where that fits in one window, a single agent reads the late contracts on a crowded desk, or has to compact as it goes, and the last contracts get a tired reader. Eight workers with five contracts each keep every desk small. Every row of the table lands on the left.
What does it cost, and what does it buy?
Assume a single agent would take about three minutes per contract, so two hours end to end. Eight parallel workers finish their five contracts in about fifteen minutes, and the orchestrator needs perhaps ten more to assemble and sanity-check the table: roughly 25 minutes instead of 120. On tokens, apply the derived ratio from earlier: expect something like three to four times the single-agent bill, less if the workers run on a cheaper tier and only the orchestrator uses the strongest model.
Is that worth it? For a once-a-quarter renewal deadline, probably yes, because the gain is speed plus a fresher reader on every contract. The two disciplines that make it work are both cheap: a structured brief per worker (contract IDs, the clause definition, the exact output schema, “do not compare contracts, that is the lead’s job”) and a citation check on every returned row.
Why does the migration stay single?
Proposal B: add a tenant identifier to every database query in the billing service, using eight workers, one per module. It sounds just as parallel.
The table disagrees almost row by row. The modules share helpers, so branches see each other’s changes. The work is writing to one codebase, and the results combine by merge. Each worker will make silent decisions (thread the identifier through a context object or pass it as a parameter?) that must agree everywhere.
Whichever worker first changes a shared helper moves the baseline for the other seven, and that is the merge wall arriving on schedule.
The design that passes review is a single agent doing the writes, in one continuous context, module by module. Parallelism still has a place on the reading side.
A read-only subagent can survey every query site first and return a list, and a fresh-context reviewer can check the finished diff. That is the book’s rule (writes stay single-threaded while intelligence fans out) applied to a real change. If the work must be split for speed, the book’s advice is to budget the merge queue before the collision, not after it. The state-machine view of an agent is a useful lens here: one writer means one well-defined state at a time.
What changes when the orchestrator is you?
When a person orchestrates several agents directly, the coordination quality improves and the bottleneck moves to your attention. You decompose, you brief, you judge the results, and the limit becomes how many finished results you can verify, understand and integrate in a day, not how many agents you can start.
Chapter 11 gathers practitioner reports on this and finds them disagreeing about the right number of parallel agents but agreeing that a ceiling exists. Past a handful, each additional stream degrades your attention to all of them, which the book calls context rot in the operator. The useful countermeasures are the ones that lower the load: a visible status board, a work-in-progress limit, and a signal that lets a task ask for attention instead of you polling every pane. Blast radius sets the sensible degree of parallelism too: many small, isolated tasks can run at once, while a change to the heart of the system wants one agent and your full attention.
Where does this verdict not apply?
The verdict assumes the costs above are binding, and sometimes they are not. If latency matters far more than tokens, if the task is a broad search where coverage is the whole point, or if a single strong agent is clearly weak on the task, the threshold for splitting drops and the decision table should be read generously.
It also ages. The research team’s own write-up notes that tasks requiring shared context “are not a good fit for multi-agent systems today,” and the coding team’s follow-up puts it bluntly: “The open problems are all communication problems”. Both are statements about current models.
As models get better at handing context to each other, some right-hand rows will soften; as single agents get stronger, the Kim study suggests the gains from splitting may shrink instead. Either way, the procedure survives: measure a single-agent baseline, then test whether delegation earns its bill.
Nothing here replaces a measured comparison on your own workload, and the coordination failure modes in the multi-agent overhead post are what to instrument before you trust one. For what a single loop does on each step before you multiply it, the agent loop explainer is the place to start.
The takeaway
Ask the book’s question before drawing any org chart: “does this work decompose into branches that can run without seeing each other?” If yes, the orchestrator–worker pattern is on the table, with tight briefs, checked returns and single-threaded writes. If no, keep one continuous context, because splitting it breaks the coherence the task depends on. As Chapter 11 puts it, “Adding agents adds hands and coverage; it does not add judgment.”
Chapter 11, “Multi-Agent Systems,” is in the full book, with the brief-writing checklist, the coordination failure field guide, the operator-attention material and the merge wall in depth. Chapter 2, with the compounding-error arithmetic, is free to read online; see the formats.
Questions readers ask
- Is a multi-agent system better than a single agent?
- Not in general. Multi-agent systems win on read-heavy work that splits into independent branches, where parallel workers add speed and coverage. On sequential or write-heavy work they often lose: one controlled study found relative changes ranging from +80.8% on decomposable financial reasoning to -70% on sequential planning, depending on how well the architecture matched the task.
- How much more does a multi-agent system cost than a single agent?
- One team that built a production research system reported that single agents used about 4 times the tokens of a chat interaction and multi-agent systems about 15 times, which, by my own division (15 / 4), works out to roughly 3 to 4 times a single agent; the team did not report that ratio. That is one point-in-time measurement on one workload; measure your own, and remember that running workers on a cheaper model tier can keep the bill from multiplying as fast as the tokens.
- What is the difference between subagents and a single agent with tools?
- A single agent with tools keeps one continuous context and calls functions that return results into it. A subagent is a full agent loop in its own context window, briefed by the main agent and returning only a distilled result. Use subagents when you want a messy subtask's context kept out of the main window; use tools when the main agent needs to see everything.
- What is the merge wall in multi-agent coding?
- The merge wall is the point where parallel agent workers, each starting from the same baseline, try to combine their changes and discover the baseline has moved so much that their work no longer applies. Steve Yegge named it in 2025. The practical responses are a serialized merge queue, deferring overlapping work, and pausing parallelism during sweeping changes.
- When should I use multiple AI agents?
- Use several agents when the subtasks cannot be listed until the input is seen, the branches are independent by nature, the information overflows one context window, each worker's return can be checked cheaply, and the results combine by assembly rather than by integration. If any of those fails, a single agent is usually the better design.
Sources
- Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford (Anthropic) (2025). How we built our multi-agent research system
- Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al. (2025). Why Do Multi-Agent LLM Systems Fail?
- Yubin Kim, Ken Gu, Chanwoo Park, et al. (2025). Towards a Science of Scaling Agent Systems
- MIT Media Lab (2026). Towards a Science of Scaling Agent Systems: When and Why Agent Systems Work (project overview)
- Walden Yan (Cognition) (2025). Don't Build Multi-Agents
- Walden Yan (Cognition) (2026). Multi-Agents: What's Actually Working
- Steve Yegge (2025). Six New Tips for Better Coding With Agents
- Erik Schluntz and Barry Zhang (Anthropic) (2024). Building effective agents