Home / Blog / Patterns and multi-agent systems / The Orchestrator Worker Pattern: Writing the Wo…

Patterns and multi-agent systems

The Orchestrator Worker Pattern: Writing the Worker's Brief

Orchestrator worker pattern: a six-field worker brief, one task briefed badly and well, and a symptom table. Copy the template, then read Chapter 11.

By Enrique Gutiérrez · Published · 27 min read

In the orchestrator worker pattern, a lead agent decides the subtasks at run time and hands each one to a worker that starts with nothing but the message it is given. What the worker can do is therefore bounded by that message, the brief. This post is about writing it.

Most multi agent orchestration guides draw the hub and the spokes, yet the best-documented failure of this design that I found, and its fix, were both text. One lab’s lead agent sent workers “simple, short instructions”, and the workers “misinterpreted the task or performed the exact same searches as other agents” (Hadfield et al., 2025). So the unit of design here is the delegation message: six fields, one task briefed three ways, and a table from symptom to fix.

Whether to split the work at all is a separate decision, and the single agent vs multi agent comparison owns it, cost argument included. Its short version: start with one agent, and split only when the branches can run without seeing each other. Everything below assumes the subagents vs single agent question is settled.

What is the orchestrator worker pattern?

The orchestrator worker pattern is a design in which one agent reads the task, writes the list of subtasks itself, hands each subtask to a worker with its own context window, and then combines what comes back. Chapter 11 of AI Agents, Engineered (in the full book) defines the roles in its section “The Orchestrator–Worker Pattern”: “An orchestrator (also called a lead or supervisor) receives the task, breaks it into subtasks, and hands each one to a worker”.

The name was fixed by a 2024 essay that lists it among its agentic workflow patterns and separates it from a fixed fan-out: the subtasks “aren’t pre-defined, but determined by the orchestrator based on the specific input” (Erik S. and Barry Zhang, 2024). That is the property this post depends on. In a fixed workflow you write the task messages once; in an orchestrator worker system a model writes every delegation, fresh, on every run.

What does a worker know when it starts?

A worker knows what its brief says and nothing else. It is a subagent with a clean context window, and Chapter 11’s section “Focused, Self-Contained Worker Context” seats you at its desk: “You did not hear the user’s request; you were not in the room when the orchestrator planned; you cannot see what your five siblings are doing, or that siblings exist at all.”

A worker inside its own context window sees only the memo the orchestrator wrote, while the user's request, the planning, and its sibling workers stay outside, unseen.
Figure 11.2 A delegated worker’s entire world: an empty desk, one memo, and siblings it can never see. The worker is exactly as good as the memo that summoned it. Reuse this diagram

Then: “On the desk in front of you is a memo, and the memo is everything.” A worker is isolate, the last of the four context moves the book teaches (write, select, compress, isolate), turned into an architecture. The price of the clean window is that nothing crosses into it except the brief.

Users describe the same desk from outside. One, on a coding agent’s subagents in a 2025 forum comment: “They don’t know what you are doing, they don’t care what you do with the info they produce.” Comments like that one are complaints, found by searching for complaints, mostly from coding-agent users. They show what the failures look like, and nothing about how often they happen.

What goes in a worker’s brief?

A worker’s brief needs six fields: an objective, an output format, tool and source guidance, boundaries, a stop condition, and an instruction for when the worker is blocked. The first four are the book’s. The last two are this post’s additions, and are labeled as such below.

The book’s checklist, from the same section: “A good delegation gives the worker four things. An objective: the one thing this worker is responsible for producing. An output format: the shape the result must come back in, so that synthesis is assembly work instead of a second round of interpretation. Tool and source guidance: which tools to use and where to look”.

The fourth is “boundaries: explicitly what this worker should leave alone, because it belongs to a sibling.” The book also supplies the test: “write every brief as if it were a ticket for a contractor who has never seen your project and cannot ask questions.”

Stop condition (this post’s addition). The worker must know what “enough” is and where the ceiling sits. The same lab’s published worker prompt, as read in October 2026, caps tool calls and adds: “when you are no longer finding new relevant information and results are not getting better, STOP using tools and instead compose your final report.” That team keeps these lines in the worker’s standing prompt; I put the stop condition in the brief, since “enough” differs by subtask.

When blocked (this post’s addition). The worker must know what to send back when it cannot finish honestly. The same prompt has one such line: “If unable to reconcile facts, include the conflicting information in your final task report for the lead researcher to resolve.” A coding-agent team wrote in 2026, about a stronger model advising a primary model that had not read a file its question depended on, that “the right answer from the smart model is not to make up some theories (which is often the default behavior)” (Yan, 2026).

Both gaps appear in a published failure taxonomy. MAST (Cemri et al., 2025, version 3 of the preprint) puts “unaware of termination conditions” at 12.4% and “fail to ask for clarification” at 6.80%, as prevalence among 14 failure modes across 1,642 annotated traces; the 14 figures sum to about 100, so I read each as a share of all labeled failures. Those traces come from seven multi-agent frameworks, so the figures describe multi-agent systems in general, not this pattern.

Field Whose What the worker reads Without it
Objective The book’s The one thing to produce, why the lead needs it, what is already known A topic gets explored; no deliverable arrives
Output format The book’s A status line, then the shape and length of the final message The lead interprets prose a second time
Tool and source guidance The book’s Which tools, where to start, what counts as reliable Careful work in the wrong place
Boundaries The book’s What a sibling owns, what is already decided, where it may write Overlap, gaps, clobbered files
Stop condition This post’s What “enough” is, and a ceiling on effort A run that never ends, or ends at the first hit
When blocked This post’s What to report when something cannot be established or the brief itself looks wrong A guess written in the register of a fact

The sub-lines inside each field (the status line, already known, already decided, where to write) are my elaboration; the book names the four fields and stops. Two of them have a published precedent: the same lab’s lead prompt, as read in October 2026, asks for “Relevant background context about the user’s question and how the subagent should contribute to the research plan” and tells the lead to “define what constitutes reliable information or high-quality sources for this task”.

The six fields overlap at the edges (a “not published” entry is both an output value and a blocked case); they are prompts for the writer, not separate compartments. I have not measured four-field against six-field briefs and found no published comparison: the two additions are an argument from the sources, not a result.

The six-field worker brief template

This template is the six fields as plain text with placeholders, in no product’s syntax. Paste it into the orchestrator’s prompt as the form it must fill for every delegation, or into the function that composes a worker’s task message. Tell the orchestrator to omit a line that does not apply to the subtask at hand, and never to leave a placeholder unfilled. The book calls its own sample brief “illustrative, not a template to copy”, and the same holds here: this is a form to fill, not text to reuse.

OBJECTIVE
Produce: <the one deliverable this worker owns>
Why the lead needs it: <one sentence>
Already known or ruled out, do not re-derive: <facts the lead has established>

OUTPUT FORMAT
First line: DONE, CEILING REACHED or BLOCKED, then what you did not cover.
Then: <headings or fields, in order>   Length: <cap>
For each claim: <where it came from: URL, file and line, or query>
Keep findings separate from interpretation.
Your final message is the only thing the lead will see.
Large artifacts: write to <location>; return the reference and a short conclusion.

TOOL AND SOURCE GUIDANCE
Use: <tools>   Start from: <sources, paths or tables>
Prefer: <what counts as reliable here>   Avoid: <what does not>

BOUNDARIES
Not yours: <what each sibling owns>
Already decided for every worker: <units, naming, formats, conventions, and the key the lead will join results on>
Write only to: <path or store>   (or: read-only)

STOP CONDITION
Done when: <the observable state that means enough>
Effort ceiling: <a number of tool calls, sources or minutes>
Stop earlier if <two attempts in a row add nothing new>.

WHEN BLOCKED
If you cannot establish something: say so, say what you tried, say what you would need.
If sources or facts conflict: report both. Do not pick one silently. Do not guess.
If finishing would take you outside your boundaries, or the brief's premise looks wrong: stop and report. Do not work around it.

One task briefed three ways: what does each worker return?

A thin brief returns the worker’s guess at the task; a six-field brief is written to return a piece the lead can assemble. Take the task Chapter 11 uses to introduce the orchestrator worker pattern: summarize the competitive field for managed vector databases, where the lead has surveyed the market and picked six vendors. Every brief and every return below is constructed for this post. They are what I would expect, not logs of observed runs.

Undivided and thin Divided but thin Six-field brief
What the lead sends “Research managed vector databases.” to three workers “Research vendor C’s managed vector database.” and five like it, to six workers The brief below, one vendor each, to six workers
What each worker decides alone Which vendors, which facts, how deep, when to stop Which facts, how deep, what shape, when to stop Which pages to read, what counts as a limit, what to drop to fit the length
What plausibly comes back Three market overviews that overlap on the best-known vendors Six profiles covering different facts at different lengths, prices with the unit unstated or converted on the worker’s own assumptions, gaps unmarked Six synopses, each opening with a status line, under the same three headings; gaps marked “not published”
What the lead does next Re-reads everything, then researches the holes itself Re-reads all six, then re-researches the cells that are missing or not comparable Checks the shape in code, assembles a table, spends its effort on the comparison

The first column fails at dividing the work as well as at briefing it. The second is the fairer comparison, since it is the sentence a lead writes once it has a vendor list. My rewrite starts from the book’s own sample for this task (an objective, a 200-word synopsis under three headings, the vendor’s documentation as the preferred source, no comparing) and fills the rest of the template. The “Large artifacts” line is omitted because there are none.

OBJECTIVE
Produce: a profile of vendor C's managed vector database.
Why the lead needs it: it is building a six-vendor comparison, so it needs
the same three facts for every vendor.
Already known, do not re-derive: vendor C sells a hosted tier.

OUTPUT FORMAT
First line: DONE, CEILING REACHED or BLOCKED, then any heading you did not cover.
Then three headings, exactly: Pricing model / Scaling limits / Index types.   Length: 200 words at most.
For each claim: one source URL.
Keep findings separate from interpretation: under a last line "Inferred", anything the pages do not state.
Your final message is the only thing the lead will see.

TOOL AND SOURCE GUIDANCE
Use: web search and page fetch.   Start from: the vendor's documentation and pricing page.
Prefer: the vendor's own pages, plus one independent source for the limits.   Avoid: forum posts.

BOUNDARIES
Not yours: vendors A, B, D, E and F. Do not compare vendors; the lead does that.
Already decided for every worker: prices in US dollars, in the unit the vendor
publishes (per month, per unit stored, per query). Do not convert. The lead joins
results on the vendor name as written in this brief.
Read-only.

STOP CONDITION
Done when: each heading holds a sourced entry or the words "not published".
Effort ceiling: 10 tool calls.
Stop earlier if two searches in a row add nothing new.

WHEN BLOCKED
If a fact is not public: write "not published" and give the page you checked.
If two sources disagree: report both. Do not estimate.
If vendor C has no managed offering, or a heading cannot be filled without
comparing vendors: stop and report. Do not work around it.

The ceiling of 10 is illustrative; what matters is that it is a number. A return that follows this brief looks like the block below, with placeholders where a real run would have facts:

DONE. All three headings covered.

Pricing model
<the vendor's published rate, in the vendor's own unit> <URL>

Scaling limits
Not published. Page checked: <URL>

Index types
<the index types the documentation lists> <URL>

Inferred
Nothing.

The same rewrite on a coding task

Code takes the same rewrite when the worker reads and reports. Suppose a lead is about to replace scattered retry logic with one shared wrapper, and first sends workers to survey the code. The thin brief, “Look into how we handle retries,” plausibly returns four paragraphs on retry strategy and no way to tell whether the survey is complete. The rewrite is constructed too:

OBJECTIVE
Produce: a list of every place the application code retries an outbound HTTP call.
Why the lead needs it: it must decide whether one shared wrapper can replace them all.
Already known, do not re-derive: the payments module uses the shared client.
List it as one row and do not open it.

OUTPUT FORMAT
First line: DONE, CEILING REACHED or BLOCKED, then the directories you did not reach.
Then a table: file / function / retry count / backoff / idempotency key sent (yes or no).
Then three lines of conclusion.   Length: the table plus those three lines.
Then a list headed "Unresolved", if any.
For each row: file path and line number.
Keep findings separate from interpretation: the table is findings; the three lines are yours.
Your final message is the only thing the lead will see.

TOOL AND SOURCE GUIDANCE
Use: code search and file reads.   Start from: the HTTP client module and its callers.
Prefer: the code on the main branch.   Avoid: comments and documentation as evidence of behavior.

BOUNDARIES
Not yours: the test directories, which a sibling is surveying.
Already decided for every worker: paths relative to the repository root. The lead
joins results on file path.
Read-only.

STOP CONDITION
Done when: a search for each of the client's call names returns no file you have not read.
Effort ceiling: 25 tool calls.
Stop earlier if two searches in a row return only files you have already read.

WHEN BLOCKED
If a call is built dynamically and you cannot tell whether it retries: list it
under "Unresolved" with file and line.
If the code and its configuration disagree on a retry count: report both.
If the survey would need a change to any file, or there is no single HTTP client
module: stop and report. Do not work around it.

The status line does the work the thin brief could not. A worker that stops at 25 calls must say so and name the directories it never reached, so a partial survey cannot pass for a complete one.

How much context should a worker get?

A worker should get what a competent stranger needs to do the job, and nothing that would not change what it does. The 2026 essay quoted above, by the author of a well-known 2025 argument against multi-agent systems, leaves the matter open: “How do you transfer context between agents without drowning the receiver?” (Yan, 2026).

That essay reports what one kind of over-briefing, scripting the method, looked like in its authors’ own system: “Managers trained on small-scoped delegation default to being overly prescriptive, which backfires when the manager lacks deep codebase context.” The lab’s published lead prompt, as read in October 2026, asks the lead to “maintain extremely high information density while being concise - describe everything needed in the fewest words possible.”

What a worker inherits from the harness is a separate cost. A bug report against one coding agent from September 2026 measured a 396-character task: a default worker’s first call, in a session with about ten tool connectors attached, carried 60,236 tokens, and a worker limited to three read-only tools carried 5,332. That is one user, one product, one day; look up what your own harness hands a worker.

For the brief itself I use four rules, which are this post’s judgment, not the book’s:

  1. Run the deletion test. Delete any sentence whose removal would not change what the worker does or returns. One sentence of purpose usually survives; the lead’s reasoning history and the user’s conversation usually do not. The test is only as good as your model of the worker; the 2025 essay’s objection is that you cannot always know which detail will matter.
  2. State the outcome, the boundaries and a starting point; do not script the steps wherever the worker will know the terrain better than the lead after one look.
  3. Name what was ruled out when the worker would otherwise rediscover it.
  4. Withhold the lead’s guesses and conclusions from any worker sent to check or establish them. They would change its behavior, and that change is the bias you are paying to avoid. The book’s reviewer archetype is “handed only the finished work and the standard to judge it against”.

Take the vendor brief with three sentences a lead might add. The example is constructed, and each sentence should be struck:

Sentence added to the brief Struck by What it would do to the worker
“I picked these six vendors because they came up most often in my survey.” Rule 1 Nothing: the profile of vendor C is the same without it, and the worker re-reads the sentence on every step
“Search the vendor’s name plus ‘pricing’, open the first three results, then search for ‘limits’.” Rule 2 The worker follows the script past the pricing page that the documentation links directly
“I believe vendor C charges per query.” Rule 4 The worker looks for confirmation of the lead’s guess and reports it as found

On the objection in rule 1: the 2025 essay holds that in a real conversation “any number of details could have consequences on the interpretation of the task” (Yan, 2025), and its remedy was one continuous context. I cannot remove that objection. A brief is a bet that the lead can name the details that matter, and the when-blocked field is for the times it loses.

How do you stop workers duplicating or contradicting each other?

Two lines in the boundaries field address the duplication a brief can address, and one line in the objective keeps the worker from repeating the lead. Name what each sibling owns, state the shared decisions in the same words in every brief, and say what is already known. The lab’s published lead prompt puts the first of these bluntly: “Avoid overlap between subagents - every subagent should have distinct, clearly separate tasks”.

A constructed data example shows what the second line is for: suppose three workers each profile one month of an orders export. Every brief carries the same sentence: amounts are integer cents, timestamps are UTC, null values are counted and never dropped, and the lead joins on order ID. Every brief also says which month is the worker’s and that removing duplicates across months is the lead’s job. Without the shared sentence the three profiles can each be correct and still not add up.

Beyond that, a brief runs out. Chapter 11’s section “Coordination Failure Modes” describes four seams. It names two, duplicated work and gaps and conflicting implicit assumptions, calls the third “the funnel itself”, and leaves the fourth unnamed; I’ll call it the trusted return. The first is the brief’s home ground: “A brief is the cheapest artifact in the whole system to improve, and the most commonly skimped.”

Good briefs do not close the second seam. The book again: “No brief can enumerate every decision in advance; briefs prescribe what you thought to prescribe.” A shared-decisions line catches the conflicts you can name; the rest need one writer, covered below.

How many workers should an orchestrator spawn?

Spawn one worker per independent subtask, under a ceiling written into the orchestrator’s own prompt. The rule is not a worker-brief field. The lab reports an early version “spawning 50 subagents for simple queries” until rules went into the prompt (Hadfield et al., 2025).

The published rules are examples of having a rule. The lab’s 2025 article says “Simple fact-finding requires just 1 agent with 3-10 tool calls, direct comparisons might need 2-4 subagents with 10-15 calls each, and complex research might use more than 10 subagents with clearly divided responsibilities.” Its published lead prompt, as read in October 2026, gives four tiers ending in “5-10 subagents (maximum 20)”.

Those two rules come from the same team and do not match. My six-vendor example lands in the prompt’s top tier and between two tiers of the article’s, so write your own tiers and date them. A tier is one line: a kind of task, a worker count and a per-worker ceiling, for example “one-per-item comparison: one worker per item, at most 8, 10 tool calls each” (illustrative).

The product of the two numbers is an upper bound you can read before the run: six workers at 10 calls is at most 60 tool calls, plus the lead’s own. For what the published token multipliers imply, see the comparison post linked at the top; I found no multiplier measured by an independent team.

For each subtask I use three conditions: delegate when it is independent of its siblings, generates bulk the lead will not need afterward, and returns something cheap to check. The first two restate a practitioner test that Chapter 11 quotes; the third is the chapter’s condition for running parallel attempts. If any of the three fails, or the lead already knows the answer, the lead does the step itself.

What should a worker send back, and can you see what was sent?

A worker should send back one final message in the shape the brief fixed, and you should be able to read both the brief and the return afterward. Chapter 11 lists three communication patterns: return-values-only, the shared scratchpad and message passing. Its advice is to “start with return-values-only, and upgrade only when an observed failure asks for it.”

Return-values-only has a cost, stated by the lab that runs it: “the lead agent can’t steer subagents, subagents can’t coordinate” (Hadfield et al., 2025). The brief is the only steering there is. Large outputs should travel by reference. The book: “However your agents talk, keep large artifacts out of the conversation”, because “The orchestrator needs to know the artifact exists, where it lives, and what it concluded.”

One template line, “Your final message is the only thing the lead will see,” is necessary and not sufficient. A bug report from August 2026 describes a worker in one coding agent that ran 65 tool calls and “returned a final message that describes a deliverable instead of containing one”, although, by the reporter’s account, an instruction saying the final message was the deliverable had already been injected into its context. The reporter’s conclusion: “Prompt-level mitigation does not close this.” Keep the line, and check the return in code: if the status line or the headings the brief fixed are missing, reject it and ask again before the lead reads it.

Then log. The book’s instruction is to “record every brief and every return, because those crossings are where the bugs live”. A 2026 forum comment asks from the user’s side: “If the orchestrator agent -> subagent prompt is encrypted then how do you even know if the orchestrator agent is doing its job?” You cannot fix a brief you cannot read, so put both documents in the run’s trace.

On trust, the book needs one sentence: “a worker’s return is tool output in the sense that matters, untrusted until checked”. Anything the worker read could have carried a prompt injection.

Who is allowed to write?

One agent writes each artifact, and every other agent reads and advises. Chapter 11 states it as “the one structural rule the field has genuinely converged on: keep writes single-threaded”, and gives the short form: “one agent owns the artifact; everyone else advises.” The 2026 essay from the skeptical camp agrees in its own words: “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions.”

In practice, in my order of preference:

  1. Workers are read-only. They return findings and the lead, or one designated writer, makes every change.
  2. Each worker writes to its own path, and the lead merges. The brief’s “Write only to” line carries this.
  3. Edits to one artifact run in sequence, each starting from the last one’s result.

This rule is structural, and the brief only transmits it. It is also the step where how to build a multi-agent system stops being a prompt question and becomes a question of who owns what.

What should you change? A symptom table

A bad run of the orchestrator worker pattern can often be read backward: find the symptom, and check the field it points to. The last column says where the fix goes: one of the six fields, “too much in the brief”, or “not a brief problem” for what is fixed in the orchestrator’s prompt, the harness or the structure. The mapping is my analysis, built from the sources above. Use the buttons to narrow it; with scripts off the whole table is still here.

Symptom in the run What went wrong Fix What to fix in the brief
Two workers return the same findings, and one area goes uncovered Each got the same sentence and no statement of what the others own One objective per worker; name each sibling’s share; run pre-flight item 3 for the area nobody owns boundaries
Siblings overwrite one file, or append to one log and lose lines No brief said where each worker may write One owner per artifact; each worker writes to its own path and the lead merges (the single-writer rule is structural; the brief only carries it) boundaries
Each piece is correct and they use different units, names or definitions A shared decision was in some briefs, or worded differently in each Put the same “already decided” sentence in every brief boundaries
A worker re-derives what the lead already knew The brief did not say what was established Fill the “already known” line objective
A worker answers a neighboring question The objective named a topic, not a deliverable State the one thing to produce and why the lead needs it objective
A worker searches the wrong place or uses the wrong tool No pointer to where the answer lives Name the tools and the starting sources tool and source guidance
Findings rest on sources you would not accept No statement of what counts as reliable Say what to prefer and what to avoid tool and source guidance
The return is a wall of prose the lead must re-interpret No return shape Fix headings, fields and a length cap; separate findings from interpretation output format
The return describes the result and does not contain it The worker was not told its final message is all the lead receives, or was told and pointed at an earlier message anyway Say so, and validate the return’s shape in code before the lead uses it; for large artifacts ask for a reference plus the conclusion output format
A partial result reads as a complete one The worker hit its ceiling and nothing asked it to say so Add the status line: done, ceiling reached or blocked, and what was not covered output format
A caveat or edge case the worker found is gone from the final answer It crossed two summaries Persist the artifact; return a reference, the conclusion and the caveats verbatim output format
A worker runs far past the point of usefulness No ceiling on effort Give a numeric ceiling and a stop-early rule stop condition
A worker stops at the first plausible hit No statement of what “enough” is Write “done when” as an observable state stop condition
A worker lacked something, guessed, and the guess reads as fact No instruction for the blocked case Add the blocked line: say so, say what was tried, say what is needed when blocked
Two sources disagreed and the return quietly picked one No instruction for conflicts Add the conflict line: report both, never resolve by guessing when blocked
A worker quietly changed something outside its share to get finished No instruction for a boundary it had to cross or a premise that was wrong Add the stop-and-report line when blocked
A worker follows the lead’s scripted steps into a dead end the lead could not see The brief prescribed the method State the outcome, the boundary and a starting point; drop the steps too much in the brief
A reviewing worker agrees with everything the lead concluded The brief carried the lead’s conclusions Withhold guesses and conclusions from any worker sent to check or establish them too much in the brief
The lead spawns a crowd of workers for a small task No scaling rule in the orchestrator’s own prompt Written tiers and a worker ceiling for the lead not a brief problem
The lead delegates a step it already had the answer to No rule for when not to delegate The three delegation conditions, in the orchestrator’s prompt not a brief problem
Each piece is sound and they embody choices nobody could have listed in advance Silent decisions between parallel writers One agent writes; the others advise not a brief problem
You cannot tell which worker introduced an error Briefs and returns were not recorded Log both for every worker, keyed by run not a brief problem
A return carries an instruction the worker read somewhere The return was treated as trusted Handle every return as untrusted input not a brief problem
The lead reworded an instruction you meant to pass verbatim The model composed text that should have been fixed Send that instruction by code, not through the model not a brief problem

How do you check a brief before it runs?

Check a brief by trying to break it from the worker’s side of the desk. The list is for a brief you have already written, or one your orchestrator produced and you pulled from the log. Every item is something to verify.

  1. Hand the brief alone to a colleague, or to a fresh model session, with no access to the conversation. They can say what to produce without asking a question.
  2. Put two sibling briefs side by side. No subtopic, file or record could be claimed by both.
  3. If every worker did exactly what its brief says and no more, the lead could assemble the answer without new research, and no part of the task is left without an owner.
  4. You have written the check for the return (status line present, headings present, under the cap, a location for each claim) before the worker runs.
  5. Name one thing the worker may fail to find. The brief says what comes back in that case.
  6. Delete each sentence in turn. Every deletion would change what the worker does or returns.
  7. If the worker is there to check or establish something, the brief contains none of the lead’s guesses or conclusions about it.
  8. The brief and its return are stored where you can read them together after the run.

The third item adapts the lab’s published lead prompt, which asks the lead to make sure the results in aggregate would let it answer the user’s question.

What can a brief not fix, and when is the pattern wrong?

A brief cannot fix silent decisions between parallel writers, a harness that alters what the worker receives, or a task that should never have been split. It is the cheapest lever in the system, and only one of them.

Coordination is partly a model limit. The MAST authors write that “Solutions focused on context or communication protocols are often insufficient” for their inter-agent misalignment category, which they say demands deeper social reasoning from agents (Cemri et al., 2025). A template improves a lead’s briefs; it does not make the lead good at guessing what a worker needs.

The orchestrator worker pattern is the wrong one in three cases. You can write the subtask list in advance, so a fixed workflow is simpler. The work is mostly writing to one shared artifact; the lab that champions the pattern says “most coding tasks involve fewer truly parallelizable tasks than research” (Hadfield et al., 2025). Or the lead already knows the answer, and delegation only adds a handoff. If your doubt is one step earlier, whether the task needs an agent at all, the Should this be an agent? decision tool asks that question.

The evidence here is thin in places. The four-field checklist comes from one lab’s account of one system, and the two added fields from that lab’s prompts, a general failure taxonomy and my reading. The worked examples are constructed, with no real log among them. The dated product reports will age, and what a worker inherits is a property of your harness in a given month.

The takeaway

If you orchestrate multiple AI agents, open the log, find the delegation your lead sent on its worst run, and read it as the worker would: alone, with no history, unable to ask. Whatever you would have had to guess is a field to add. In the orchestrator worker pattern that message is the interface, and the book’s phrase for it fits: “the brief was already the interface.”

If you are choosing among multi-agent systems books to go further, look first at whether a candidate says anything about the delegation message; one that only draws the topology will leave you where the diagram did.

The delegation checklist, the three communication patterns and the four seams are in Chapter 11, “Multi-Agent Systems”, in the full book. The free guide to agent patterns collects the related posts and tools, and you can see the formats.

Questions readers ask

What should an orchestrator tell a subagent?
Six things: the one deliverable the worker owns and why the lead needs it; the exact shape of the return, opening with a status line; which tools and sources to use; what belongs to a sibling and what is already decided; when to stop, with a numeric effort ceiling; and what to send back when it is blocked or sources conflict. The four field names (objective, output format, tool and source guidance, boundaries) are the checklist in Chapter 11 of AI Agents, Engineered; the stop condition, the when-blocked instruction and details such as "why the lead needs it" and "what is already decided" are this post's additions.
How much context should a worker agent get?
Enough for a competent stranger to do the job from the text alone, and no more. Delete any sentence whose removal would not change what the worker does or returns. One sentence of purpose and a line of what is already known usually survive; the lead's reasoning history, a step-by-step script and the lead's guess at the answer usually should not.
Why do my subagents duplicate each other's work?
Often because they received the same vague sentence and no statement of what their siblings own. One lab reported exactly this in 2025: given a short instruction, one worker explored one period while two others duplicated work on another. The fix a brief can offer is one objective per worker and a boundaries field that names each sibling's share.
How do I see what my orchestrator actually sent to a worker?
Log it. Store every brief and every return, keyed by run and worker, where you can read them side by side after the run. Chapter 11 gives the reason: those crossings are where the bugs live. If your harness does not expose the delegation message, that is the first thing to fix or to ask its maintainers for.
Can workers write to the same files?
Avoid it. The rule Chapter 11 calls the one the field has converged on is to keep writes single-threaded. In practice: workers are read-only and return findings, or each worker writes to its own path and the lead merges, or the writing is done in sequence by one agent.

Sources

  1. Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford (Anthropic) (2025). How we built our multi-agent research system
  2. Anthropic cookbook (2026). Research lead agent prompt (published prompt file, undated; read 6 October 2026)
  3. Anthropic cookbook (2026). Research subagent prompt (published prompt file, undated; read 6 October 2026)
  4. Erik S. and Barry Zhang (Anthropic) (2024). Building Effective Agents
  5. Walden Yan (Cognition) (2025). Don't Build Multi-Agents
  6. Walden Yan (Cognition) (2026). Multi-Agents: What's Actually Working
  7. Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al. (2025). Why Do Multi-Agent LLM Systems Fail? (preprint, version 3)
  8. Hacker News commenter prash2488 (2025). Forum comment on subagents as freelance contractors (September 2025)
  9. Hacker News commenter kingstnap (2026). Forum comment on delegation messages that cannot be read (July 2026)
  10. anthropics/claude-code issue tracker (2026). Issue #88580: subagent returns a description of its deliverable instead of the deliverable (reported August 2026)
  11. anthropics/claude-code issue tracker (2026). Issue #97076: default subagents spawn with about 60k tokens of inherited context (reported September 2026)