Home / Blog / Patterns and multi-agent systems / How to Build a Multi Agent System That Survives…

Patterns and multi-agent systems

How to Build a Multi Agent System That Survives the Merge

How to build a multi agent system as a state machine: filled worker briefs, a merge that flags conflicts and escalates, and caps. Copy the pseudocode.

By Enrique Gutiérrez · Published · 19 min read

How to build a multi agent system, in one paragraph: write the orchestrator as a state machine in ordinary code. It plans, briefs workers, waits with a deadline, validates every return, merges by a rule chosen in advance, and sends each conflict to adjudication or a person. Caps on waves, adjudications, tokens and time end every run with a named reason.

I wrote this for a backend engineer whose split has already been approved. Someone decided that one agent drowns on the task, a framework demo looked plausible, and now the control flow is yours. Most tutorials stop at the diagram of roles. This one ends at the merge, because the merge is the step that decides whether the answer you ship is the answer your workers found.

Where does a multi-agent system fail, and why design the merge first?

Failures are mostly authored at the brief and in silent decisions, and they surface at the merge. A deliberate merge step is therefore your detector. Without one, the conflicts between workers are absorbed into a confident final paragraph and never reach a log.

The book is careful about where the damage starts. Chapter 11 of AI Agents, Engineered (Chapter 11, “Multi-Agent Systems”, in the full book) says the quality of the briefings is “where multi-agent systems actually live or die”. The glossary entry for orchestrator-worker puts it as two doors: “Nearly every coordination failure traces to a violation at one of the two doors—the brief in, the return out”.

The merge is where those failures become visible and expensive. Even good briefs leave gaps, because “No brief can enumerate every decision in advance; briefs prescribe what you thought to prescribe.” The same chapter adds: “The conflicts surface at integration, which is the most expensive possible moment to discover them”. And it calls synthesis the step “where the final answer’s quality is won or lost”.

Readers searching for how to build a multi agent system arrive at the merge with the same question. A February 2026 Hacker News comment asked how to “combine the output of 1000 subagents into one output”, adding “i think it’s a nontrivial problem”. The first reply was “You just pipe it to another agent to do the reduce step”. That reply describes the silent merge this post is built to avoid.

The measured evidence supports care at this step without making it the whole story. In the third version of the MAST failure taxonomy (Cemri et al., arXiv, revised October 2025), the three verification modes together are 23.5% of the failure labels (6.20 + 8.20 + 9.10, my sum). That is a share of labels assigned to benchmark traces. It does not say how often multi-agent systems fail, and the overhead numbers belong to the multi-agent overhead post.

What is the outer loop of an orchestrator?

The outer loop of an orchestrator is decompose, gather, assess, and perhaps decompose again. Chapter 11 states it in one sentence: “That loop—decompose, gather, assess, perhaps decompose again—is what lifts the pattern above a one-shot fan-out.”

Assessing means answering a specific question: “is this enough to synthesize, or did what came back reveal a gap worth a second, smaller wave?” One production account says the same in a builder’s words. “The LeadResearcher synthesizes these results and decides whether more research is needed—if so, it can create additional subagents or refine its strategy” (Hadfield et al., 13 June 2025).

Chapter 13, “Writing the Outer Loop” (in the full book) supplies the other half. Its rule is short: “Pair every goal with a budget.” The chapter’s reason is that “the goal function and its limits are designed together, or the loop is not designed at all”. It also links the two chapters: “a loop-launched agent is a worker whose orchestrator happens to be a script”.

The outer loop of a single agent, one item per pass, is the subject of the AI agent as a state machine. What follows reuses its four terminal states and their reasons, so the two designs agree. It adds only what fan-out needs: many returns at once, a deadline, a merge, conflicts and waves.

How to build a multi agent system as a state machine?

Write nine working states and four terminal ones, and let code move between them. The model plans, briefs and adjudicates inside states; it never chooses the transition. Every caption below is my construction on the book’s verbs, and every number in it is illustrative.

ORCHESTRATOR, one task. Caps are illustrative: set your own.
  waves <= 2 per run · adjudications <= 3 per run · wave deadline 10 min
  token and wall-clock ceilings per run

STATE          WHAT HAPPENS                                       NEXT
PLANNING       read the task and the record; write the subtask    BRIEFING
               list (one owner each), the join key, the return
               schema and the merge kind; log the plan
               branches must see each other's choices          -> ESCALATED, reason "one context"
BRIEFING       fill one brief per subtask (wave 2: gaps only);    DISPATCHED
               log each brief verbatim
DISPATCHED     launch workers, each isolated and read-only,       WAITING
               or writing only to its own path
WAITING        collect returns until all are in or the wave       VALIDATING
               deadline passes; a late worker is marked
               MISSING and is not awaited
VALIDATING     check each return in code: schema, status line,    MERGING
               keys inside the worker's boundary, evidence on
               every row, size under the return cap; a bad
               shape gets one re-ask inside the deadline,
               then MISSING
MERGING        run merge(): rows, CONFLICTS, GAPS, flagged        ADJUDICATING if CONFLICTS,
                                                                  else CHECKING
ADJUDICATING   per conflict, while the cap lasts: a fresh,        CHECKING
               read-only worker sees both claims and both
               evidence trails and returns A, B, BOTH_SCOPED
               or UNRESOLVED with its reason; code writes the
               verdict into the rows; past the cap, or once the
               token or time ceiling is reached, a conflict
               stays UNRESOLVED
CHECKING       a checker that did not merge: every subtask has    DECIDING
               rows or is a gap; every row traces to evidence;
               a row that fails, or is still unknown, adds its
               subtask to GAPS; list what is still UNRESOLVED
DECIDING       first rule that fits, in this order:
               1 check passes, no gaps, nothing unresolved     -> DONE
               2 anything UNRESOLVED                           -> ESCALATED, reason "needs a person"
               3 from wave 2 on, this wave added no rows
                 and closed no gaps                            -> STUCK, reason "no progress"
               4 a wave, token or time budget is spent         -> OUT_OF_BUDGET, which cap
               5 gaps remain                                   -> BRIEFING, next wave, gaps only

Terminal: DONE, ESCALATED, STUCK, OUT_OF_BUDGET. Every report carries
the rows, CONFLICTS, GAPS, flagged rows and the reason code.
Never: WAITING -> DONE, MERGING -> DONE. A worker's "done" is a note.
Log every transition with the brief, return or verdict that caused it.

Why does each state exist?

Each state answers a failure someone has already reported. WAITING has a deadline because the production account above admits that “the entire system can be blocked while waiting for a single subagent to finish searching”. A late worker becomes a gap that the next wave can fill.

VALIDATING exists because a return is evidence to be checked. The book’s phrase is “untrusted until checked”. The MAST authors make the matching point about where checks sit: “sole reliance on final-stage, low-level checks is inadequate”. So each return is checked on arrival and the merged result is checked again by a separate state, which follows Chapter 13’s rule to “separate the checker from the maker”.

ADJUDICATING uses a fresh, read-only worker with a cap. That follows Chapter 11’s advice on reviewer agents: keep them read-only “and cap the loop, because generator and critic can circle each other indefinitely”. The cap matters because orchestrators find ways around soft limits. A July 2026 forum comment described a verification step that “turned into an attempt to throw a party with 41 … verifiers. It will find a way.”

DECIDING mirrors Chapter 13’s decide step: “Done means halt; blocked means escalate to a human; otherwise pick the next item and go again.” The order of its rules is fixed, so two engineers reading the same log reach the same verdict. Delegation loops are an old complaint: one framework’s issue, opened in March 2024, is titled “allow_delegation=True leading to infinite loop” and kept drawing reports of the same loop after a bot closed it as stale.

What does a crash do to this machine?

A crash in the middle of WAITING should cost you the unfinished workers only. Finished returns should be read from a record on restart. That is the territory of durable execution for AI agents, where a runtime replays recorded results instead of recomputing them. The glossary’s entry on durable execution puts it in seven words: “The code replays; the world does not”.

A worker that is dispatched again must not repeat a side effect, which is the subject of idempotent tools and safe retries. Read-only workers make this easy. That is one more reason to keep them read-only.

What do two filled worker briefs look like?

They look like the six-field form in the orchestrator worker pattern post, filled for one real split, with the shared decisions written identically in both. I use a constructed task a backend team meets often: before removing a deprecated field, customer.legacy_tier, from an internal API, find every consumer.

Worker A surveys the long-running services. Worker B surveys analytics jobs and scheduled exports. Both return rows keyed on the same consumer identifier, which gives the merge something to join and to compare.

BRIEF A · services
OBJECTIVE      List every service that reads customer.legacy_tier, so the
               field can be removed safely. Known: the API gateway forwards
               the field unchanged; skip it.
OUTPUT         First line: DONE, CEILING REACHED or BLOCKED, then what you skipped.
               One row per consumer, at most 2,000 tokens in all:
               consumer_id | reads_field (yes / no / unknown) | evidence | confidence
SOURCES        Code search over the service repositories and each service's
               deployed configuration. A yes cites file and line. A no names
               where you looked.
BOUNDARIES     Yours: long-running services. Not yours: analytics jobs and
               scheduled exports (worker B).
               Shared: consumer_id is the deploy name. "Reads" means any code
               path in that deploy that loads the field, serializers and SQL
               included. Unknown is allowed; no needs evidence. Read-only.
STOP           Done when every service in the inventory has a row.
               Ceiling: 40 tool calls.
WHEN BLOCKED   Mark the row unknown and say what you would need. If two
               sources disagree, report both.
BRIEF B · jobs and exports
OBJECTIVE      List every analytics job and scheduled export that reads
               customer.legacy_tier, so the field can be removed safely.
OUTPUT         First line: DONE, CEILING REACHED or BLOCKED, then what you skipped.
               One row per consumer, at most 2,000 tokens in all:
               consumer_id | reads_field (yes / no / unknown) | evidence | confidence
SOURCES        The scheduler's job definitions and the warehouse's saved
               queries. A yes cites the job file and line or the query id.
BOUNDARIES     Yours: jobs and exports. Not yours: long-running services
               (worker A).
               Shared: consumer_id is the deploy name. "Reads" means any code
               path in that deploy that loads the field, serializers and SQL
               included. Unknown is allowed; no needs evidence. Read-only.
STOP           Done when every scheduled job has a row. Ceiling: 40 tool calls.
WHEN BLOCKED   Mark the row unknown and say what you would need. If two
               sources disagree, report both.

The shared paragraph does most of the work at merge time. A deploy can be both a service and the owner of a nightly export, so both workers may see it. Because both briefs define “reads” the same way, a disagreement on that deploy is a disagreement about facts.

How should the merge step detect conflicts?

The merge should be a function written before the first run, with one declared kind, a join key and a schema, and it should return conflicts as data. It decides nothing that the evidence has not decided. Chapter 11 wants each return shaped “so that synthesis is assembly work instead of a second round of interpretation”, and the merge is where that pays.

These four kinds are my construction. They follow the book’s contrast between work whose “findings compose by juxtaposition, because a report tolerates seams” and work where “results compose by merge; seams are defects”.

merge(returns, kind, key, schema)        # code; no model call inside
  valid   = returns that passed VALIDATING
  GAPS    = subtasks with no valid return
  flagged = rows with no evidence, rows outside the worker's
            boundary, rows marked unknown     # reported, never dropped

  ASSEMBLE    disjoint pieces expected; join on key; first rule that fits
    one return has the key            -> keep the row, cite it
    several returns, values agree     -> keep one row, cite all
    known values agree, rest unknown  -> keep the known row, flag the unknown
    known values differ               -> CONFLICTS += key, each claim and
                                         its evidence; keep no value yet
  VOTE        k workers answered the same question
    a strict majority agrees          -> that answer, with the tally
    otherwise                         -> CONFLICTS += every candidate
  SELECT      k attempts at one artifact
    keep the best by a check outside the makers (tests, schema, citations)
    no attempt passes                 -> GAPS += the subtask
  INTEGRATE   several edits to one artifact
    never merge in parallel: queue the edits through one writer,
    one at a time, each starting from the last result

  return rows, CONFLICTS, GAPS, flagged

Budget, illustrative, set your own:
  adjudications <= 3 per run    past the cap a conflict stays UNRESOLVED
  waves <= 2 per run            a later wave re-briefs gaps only
  merge input <= workers x return cap (here 2 x 2,000 tokens); a return
    over its cap is rejected in VALIDATING and never truncated here

Why escalate instead of letting the merger pick?

Because a pick hides the disagreement, and the disagreement is often the finding. A February 2026 comment on a shared file store for AI agents put the problem exactly: “With prose, two agents can produce logically conflicting conclusions without triggering a merge conflict.” The tool’s author replied that merging both “is worse than surfacing the conflict”.

Chapter 11 describes what happens otherwise: “the lead ends up reconciling accounts of work rather than work”. The adjudicator in my state machine sees both evidence trails, which is the work. When the evidence cannot settle it, a person gets the conflict in the report.

Two frameworks, read on 7 October 2026 and named only as examples of graph frameworks, make the same point in their design. One raises an error when parallel branches write one state key in the same step, until you “define a reducer that combines multiple values” (LangGraph documentation). The other’s example stores each parallel branch’s result under its own key, and the page says “If you need communication or data sharing between these agents, you must implement it explicitly” (Google Agent Development Kit documentation). The transferable idea is that a merge is declared per key.

Where do the four merge kinds come from?

Each kind has a source. Steve Yegge, relaying an observation of Gene Kim’s, wrote in December 2025 that “the workstreams have a monoidal shape, and you can merge their work by doing things like summing counts or whatever”; that is assemble. Voting has support from a study reporting that “simply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiated” (Li et al., TMLR, 2024; I read the abstract).

Select is the book’s parallel-attempts case: several tries, judged by a check none of them wrote. Integrate is the merge wall, the subject of the figure below.

A task fans out to four parallel workers, but their results must pass one at a time through a single merge-queue gate in a wall to become one artifact again.
Figure 11.5 Fan-out is cheap and orderly; the reduce is a wall with one narrow gate. Parallel branches must serialize into single-file writes to become one artifact again. Reuse this diagram

For integrate, Yegge’s own remedy is a queue: “You need to serialize the rebases, and give each worker enough context, and context-window space, to fully merge their work into the new baseline.” He is candid that the step “is often messy and not entirely automatable”. Chapter 11’s verdict asks you to expect the wall where branches meet one artifact, “budgeting for the queue before rather than after the collision”.

Does the merge step hold on three cases?

It does, and each case ends in one outcome. I walked the worked example through the two pseudocode blocks above, with the illustrative caps.

Case What the workers return What merge() does Where the run ends
Two workers agree A and B both report invoice-renderer reads the field, each citing its own file and line Keeps one row, cites both CHECKING passes, then DONE
Two workers disagree on a fact A reports ledger-sync does not read the field, citing the service code. B reports it does, citing the deploy’s nightly export query Adds a conflict with both claims and keeps no value ADJUDICATING: the fresh worker returns B, because the shared definition counts SQL in the deploy. Row is yes, with both trails and the verdict logged, then DONE. Had it returned UNRESOLVED: ESCALATED
One worker times out B misses the wave deadline and is marked MISSING B’s subtask becomes a gap; A’s rows merge Rule 5: wave 2 re-briefs B’s subtask only. If B returns, DONE. If B times out again, the wave added nothing, so rule 3 gives STUCK

The third case shows why the order of rules is written down. A second timeout could read as an exhausted wave budget or as no progress. Rule 3 comes before rule 4, so the report says “no progress”, which tells the person to look at B’s sources before buying another run.

How likely is it that every worker returns something usable?

Less likely than intuition says, because every worker must succeed for the merge to complete. With four workers that each return a usable result 90% of the time, all four do so 0.94 = 65.6% of the time, assuming they fail independently. These inputs are illustrative.

A re-ask helps only if code catches the bad return; for a worker that times out, the next wave plays the role of the re-ask. With one re-ask after a failed check in VALIDATING, each worker succeeds 1 − 0.12 = 99% of the time, and all four 0.994 = 96.1%. To reach 90% across four workers without the re-ask, each would need 0.9^(1/4) = 97.4%.

The compounding error calculator speaks of steps; read each step as one worker’s return. The arithmetic is the one Chapter 2 develops for a single agent’s steps, and Chapter 2 is free to read. The calculator below opens on these numbers.

With JavaScript on, the Compounding error calculator runs here, filled in with the example from this post.

Runs in your browser; nothing is sent anywhere. Open the Compounding error calculator on its own page to share a result by link.

Independence is the optimistic case. Workers that share a misleading source fail together, and a re-ask with the same brief tends to fail the same way.

How do you orchestrate multiple AI agents without a standing integrator?

Keep the merge a function plus a capped adjudication, and give no agent the permanent job of resolving conflicts. One team’s experience is the strongest evidence for this choice. Wilson Lin wrote for Cursor on 14 January 2026: “We initially built an integrator role for quality control and conflict resolution, but found it created more bottlenecks than it solved. Workers were already capable of handling conflicts themselves.”

I read that as counter-evidence to any design that routes every return through one more agent. The same post warns from both sides: “Too little structure and agents conflict, duplicate work, and drift. Too much structure creates fragility.” My design keeps code at the center and calls a model only for the conflicts the code found.

The Cursor system also had a step that resembles DECIDING: “At the end of each cycle, a judge agent determined whether to continue, then the next iteration would start fresh.” In my version, that judgment is a rule over counters in the record, so the model has no vote on whether to continue.

How do you know you should not have built it?

Your own logs tell you, within a few weeks of runs. The pre-build question, whether to split at all, is the job of single agent vs multi agent. The list below is for a system that already runs, and every threshold in it is my illustration.

  • Most runs end in one wave with no conflicts, and a single-agent baseline on the same tasks scores the same. The split buys only cost (illustrative: three runs in four, over twenty or more runs).
  • The same keys arrive from several workers on most runs. The subtasks are not independent, and Chapter 11’s deciding question has failed: branches that “run without seeing each other”.
  • Adjudications hit their cap on more than one run in five, or most conflicts are about style, units or interpretation. Silent shared decisions are leaking between workers; one context should own the work.
  • The merge input is larger than the work it summarizes, or workers re-read the same sources. Coordination costs more than the work. One vendor’s January 2026 write-up reports a coding split whose subagents “spent more tokens on coordination than on actual work” (Phillips et al., 2026).
  • Two workers write to one artifact and the edit queue stalls. That is the merge wall; move to one writer.
  • You cannot replay a run from its log of briefs, returns and merge decisions. Chapter 11’s instruction is to “record every brief and every return, because those crossings are where the bugs live”.
  • The subtask list is the same for every input. It was a workflow all along, and a fixed pipeline is cheaper to run and to test.

Each item can be computed from the trace of a run, provided the orchestrator logs its transitions as the state machine says. Phillips et al. give no method or sample for their figures, so treat that line as one team’s report.

Where does this design stop helping?

It stops where conflicts cannot be expressed as fields. The merge detects disagreement by key and by value, so its detector is only as good as the return schema. Free-text findings need claims pulled into fields first, and that extraction is a model step with its own error rate.

The adjudicator is still a model. A fresh window and two evidence trails make a good verdict likely, and they do not make it certain, so the verdict and its reasons go in the log for a person to sample. Cognition’s April 2026 essay is blunt about the remaining gap: “The open problems are all communication problems” (Yan, 2026).

Escalation has a cost of its own. A person who receives three conflicts per run will soon stop reading them, so the adjudication cap and the escalation rate belong on the same dashboard. And the caps here are illustrative; the book gives you the rule to have them, and your own runs give you the numbers.

The literature on how agents coordinate is older than language models. Contract nets and blackboards are surveyed in multi-agent systems books, which is the reading list I would give someone who wants the theory behind the shared scratchpad. One controlled study’s abstract adds a caution on topology: “architectures without centralized verification tend to propagate errors more than those with centralized coordination” (Kim et al., 2026).

What to build this week

That is how to build a multi agent system that survives the merge: decide the merge kind and the join key before the first run, write the merge as code that returns conflicts, and cap every loop in the orchestrator. Then walk your own three cases, agree, disagree and time out, through it and make sure each one ends in exactly one state.

The worker’s brief is covered in Chapter 11, “Multi-Agent Systems”, together with the seams and the merge wall, and the decide step and its budgets in Chapter 13, “Writing the Outer Loop”. Both chapters are in the full book; the Preface, Chapters 1 and 2 and the glossary entries for budget and subagent are free. The guide to agent patterns collects the neighboring posts, and you can see the formats.

Questions readers ask

How do you combine the outputs of multiple AI agents?
Decide the kind of merge before the first run. Disjoint pieces are assembled by joining on a key fixed in every brief; answers to one question are voted on; attempts at one artifact are compared by an outside check such as tests; edits to one shared artifact go through a single writer in order. Write the merge as code, and give each worker a return schema so the merge has fields to compare.
What should happen when two agents return conflicting results?
The merge should record the conflict with both claims and both evidence trails and keep no value for that key. A fresh, read-only worker can then judge the evidence and return one claim, both claims with their scopes, or unresolved. Unresolved conflicts, and any past a fixed cap, go to a person with the report. A final model call that writes one confident answer over the disagreement hides it.
How do you stop an orchestrator agent from looping or spawning too many subagents?
Put the caps in the orchestrator's code, where the model cannot negotiate them: a deadline per wave, a maximum number of waves, a maximum number of adjudications per run, and token and wall-clock ceilings. When a cap fires, the open work is recorded: a late worker becomes a gap, a conflict past the cap goes to a person, and a spent wave, token or time budget ends the run in a named state with that work listed: out of budget, or escalated when conflicts are still open. A wave that adds nothing new should end the run as stuck rather than launch another.
Should agents talk to each other or report to an orchestrator?
Start with workers that report only to the orchestrator and never to each other. Chapter 11 of the book advises starting with return values only and upgrading when an observed failure asks for it. Shared decisions such as the join key, units and definitions belong in every brief, written the same way, so workers need not negotiate them at run time.
How do I know my multi-agent system is not worth it?
Run a single-agent baseline on the same tasks and compare four counters your logs already hold: cost per task, conflicts per merge, adjudications per run and waves per run. If most runs need one wave, find no conflicts and score the same as the baseline, the split is buying only cost. If the same keys keep arriving from several workers, the subtasks were not independent.

Sources

  1. Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford (Anthropic) (2025). How we built our multi-agent research system (engineering blog, 13 June 2025)
  2. Wilson Lin (Cursor) (2026). Scaling long-running autonomous coding (vendor blog, 14 January 2026)
  3. Steve Yegge (2025). Six New Tips for Better Coding With Agents (7 December 2025)
  4. Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, et al. (2025). Why Do Multi-Agent LLM Systems Fail? (arXiv:2503.13657, v3 revised 26 October 2025)
  5. Yubin Kim et al. (2026). Towards a Science of Scaling Agent Systems (arXiv:2512.08296, v3 revised 8 April 2026; abstract read)
  6. Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye (2024). More Agents Is All You Need (Transactions on Machine Learning Research; abstract read)
  7. Walden Yan (Cognition) (2025). Don't Build Multi-Agents (12 June 2025)
  8. Walden Yan (Cognition) (2026). Multi-Agents: What's Actually Working (22 April 2026)
  9. Cara Phillips et al. (2026). Building multi-agent systems: when and how to use them (vendor blog, 23 January 2026)
  10. LangGraph documentation (2026). INVALID_CONCURRENT_GRAPH_UPDATE (framework documentation, one example of a graph framework; read 7 October 2026)
  11. Google Agent Development Kit documentation (2026). Parallel agents (framework documentation, a second example; read 7 October 2026)
  12. amabito and syumpx (Hacker News) (2026). Hacker News comment on merging prose from parallel agents, with the tool author's reply (19 February 2026)
  13. Davidzheng and mattlondon (Hacker News) (2026). Hacker News question on combining the output of many subagents (12 February 2026)
  14. nvch (Hacker News) (2026). Hacker News comment on a verification step that spawned 41 verifiers (12 July 2026)
  15. crewAIInc/crewAI issue tracker (2024). Issue #330: allow_delegation=True leading to infinite loop (opened 8 March 2024)