Home › Study guides › CCAR-P › Domain 1 › Lesson 1.4
CCAR-P · Domain 1 · 17% of the exam · Lesson 1.4 · 22 min read
Multi-agent systems: when to split the work and how to orchestrate it
When several Claude agents beat one, how a lead agent briefs, budgets and merges its workers, and the orchestration choices an architect must justify.
Written against objective 1.4 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
1.4.1 Why one agent struggles with fifteen competitors
Every Monday the strategy team at Tesserine Pharma wants one document: what fifteen competitors did last week, from new trials and enrolment changes to partnerships and conference data. Today two analysts spend Thursday and Friday assembling it by hand from public trial registries, company press releases and conference abstracts. Declan, who runs competitive strategy, wants it automated, cited and ready before the Monday meeting. Ines, the solution architect, takes the engagement.
Her first prototype, one agent with every tool, works for three competitors. At fifteen it fails in three ways. Registry records are long, so by competitor nine the context holds tens of thousands of tokens it no longer needs, and the agent starts attaching trials to the wrong company. It searches one competitor after another, so a run takes hours. And with registry, news and abstract tools side by side, it sometimes searches the news wire for something only the registry holds.
None of this is the model's fault. One context window and one sequence of steps are holding work that is naturally separate. A multi-agent system splits it: a lead agent (the orchestrator) plans the work and hands bounded pieces to subagents (the workers). Each worker is a full agent, calling tools in a loop until its piece is done, but with its own context window and only the tools its piece needs. It returns a compact result for the lead to merge. Anthropic built the Research feature in Claude this way.
The weekly brief as a lead agent and its workers
1.4.2 When more agents help, and when they hurt
It is tempting to treat multi-agent as the advanced version of an agent, the design you graduate to. Resist it. Anthropic reports teams spending months on elaborate multi-agent architectures, only to find that better prompting on a single agent matched them. Its guidance names three situations where several agents consistently beat one; outside them, coordination costs usually exceed the benefit.
The first is breadth that parallelises: many directions that do not depend on each other. On an internal research eval, Anthropic's multi-agent system beat a single agent by 90.2%, mainly because separate context windows let it spend more tokens on the problem. The second is context isolation: a subtask that reads far more than it returns, so a worker filters it and hands back a summary. The third is specialisation: distinct tool sets, or instructions that conflict, since an agent juggling tools from unrelated domains picks the wrong one more often.
The costs are just as concrete. In Anthropic's testing, multi-agent systems typically use 3 to 10 times the tokens of a single agent on the same task. Total time is often longer too, because there is so much more computation; the real gain is thoroughness, not speed. Debugging gets harder, because runs are non-deterministic and a small change to the lead's prompt can change how every worker behaves.
| Signal in the requirement | Split or keep together | At Tesserine |
|---|---|---|
| Many items that do not need each other's findings | Split: workers run in parallel | 15 competitors, searched independently |
| A subtask reads far more than it returns | Split: the worker filters, the lead gets a summary | Long registry records become a few lines of changes |
| Tools from unrelated domains, or conflicting instructions | Split: one tool set per worker | Registry search, press wire, abstracts database |
| Steps that need each other's full context, or shared state | Keep in one agent | The cross-competitor synthesis stays with the lead |
Use the four rows as a test before drawing a second agent. Try the cheaper single-agent fixes first: tool search (loading tool definitions only when needed) for a long tool list, compaction (summarising older context) for a context that keeps filling. Neither would make Tesserine's hours-long serial run parallel. Ines justifies the token bill as Anthropic frames it: multi-agent designs suit tasks valuable enough to pay for the extra performance, and this brief replaces two analyst-days.
1.4.3 Orchestrator-worker and the alternatives
Once the split is justified, the next question is its shape. The default is orchestrator-worker: a lead decides how to approach the task, dispatches bounded subtasks, and synthesises what comes back. Each worker finishes its piece and ends. Think of an editor-in-chief assigning reporters to beats: each reporter files a story, not their raw notes, and the editor writes the front page. Precisely, the lead holds the goal and the plan; workers hold only their piece.
The pattern has two known weaknesses. The lead is an information bottleneck: anything one worker finds that another needs must travel back through the lead, and details get summarised away. And workers dispatched one after another pay the multi-agent token bill without the speed. Other shapes suit other structures of work.
| Option | When it wins | What it costs |
|---|---|---|
| Single agent with good tools | Coupled work, shared context, modest breadth | Bounded by one context window and one sequence of steps |
| Orchestrator-worker | Bounded subtasks with little interdependence; the lead must adapt the plan | The lead is a bottleneck; several times the tokens |
| Pipeline of specialists | Fixed stages, each consuming the last one's output | A failed stage blocks the rest; context thins at each stage |
| Generator-verifier | Quality-critical output with explicit criteria | The verifier is only as good as its criteria; loops need a cap |
| Handoff | The task changes owner midway, to a specialist or a person | The first agent steps out, so its context must travel with the task |
| Agent teams | Independent work that needs sustained context over many steps | Harder completion detection; teammates can conflict on shared resources |
Memorise when each row wins. Recognise two further patterns Anthropic describes: a message bus for event-driven pipelines, and shared state for agents that build on each other's findings. Anthropic's advice is to start with orchestrator-worker, watch where it struggles, and evolve from there.
Production systems combine patterns, and Tesserine's does. Collection is orchestrator-worker, with registry, press and abstract scouts per competitor. After synthesis, a pipeline stage ties every claim in the draft to its source, as the dedicated citation agent does in Anthropic's Research system. A material finding, such as a competitor halting a phase 3 trial, is handed off to an analyst before it is published.
1.4.4 Delegation: the brief, the return and the merge
Here is the failure that trips up most first designs. Anthropic's Research lead first sent short instructions such as "research the semiconductor shortage", and subagents misread the task, left gaps, and duplicated each other's searches. The fix was the brief: every delegation states an objective, an output format, the tools and sources to use, and clear task boundaries. Add an effort limit, because agents judge effort poorly on their own.
The brief has to be complete because the worker starts with none of the lead's context. In the Claude Agent SDK and Claude Code, the lead passes a subagent only the prompt string of its delegation call. Unless the subagent is a fork of the whole conversation, it never sees the lead's conversation, tool results or system prompt. It is a contractor who never attended your meetings: the work order is the whole job. Here is one of Tesserine's 45 weekly briefs; note the out-of-scope line, which prevents overlap, and the rule for failures.
Objective: find every change to Orvane Bio's trial registrations from 21 to 27 September: new trials, status changes, enrolment or end-date changes, results postings.
Out of scope: press releases and conference abstracts (other scouts cover them), other competitors, anything before 21 September.
Tools: the registry's search and get_record tools only. No web search.
Output: a JSON list, one item per change: registry identifier, programme, change type, old value, new value, date, record URL. If nothing changed, return an empty list and say so. If a tool fails, return status "failed" with the error, never a guess.
Effort: at most 15 tool calls. Known Orvane Bio trials from last week: [list]. Stop when each is checked and one search for new registrations is done.
Part of that brief never changes from week to week, so it can live in the subagent's definition. In Claude Code a subagent is a Markdown file in .claude/agents/. Its frontmatter sets name, tools, model, maxTurns and a description that Claude matches tasks against when deciding to delegate; the body becomes its system prompt. The Agent SDK takes the same fields in code, which is how Ines turned her Claude Code prototype into a service. The output schema and failure rule live in the definition; the dates, the competitor and last week's trials travel in each brief.
The return matters as much as the brief. Only the worker's final message reaches the lead, so ask for compact, structured output with a source on every item. For a large result, Anthropic suggests the worker store it and pass back a reference, avoiding a lossy game of telephone through the lead.
Structure also makes the merge mostly mechanical, and merging is where duplicates meet. Orvane's press release announcing a phase 3 start and the new registry record are one event, found twice. The lead merges on a stable key (competitor plus registry identifier) and keeps both citations. It flags conflicts instead of choosing silently: if the press release says enrolment is complete and the registry still says recruiting, the brief shows both.
One delegation, from plan to cited brief
1.4.5 Orchestration strategy: who decides the shape of the run
The next decision is who controls the fan-out. In static orchestration, your code decides how many workers run and on what. In dynamic orchestration, the lead decides at run time. Dynamic wins when you cannot predict the pieces; static wins when you can.
In Tesserine's first version the lead decided everything: in week two it skipped conference abstracts for three competitors because their press releases "looked sufficient", and later it spawned thirty extra workers chasing one rumour. Coverage of all 45 competitor-and-source cells is a requirement, so Ines moved the guarantee into code. The lead still writes the briefs, since it knows last week's context, but code rejects a plan or a merge that misses a cell. After the sweep, the lead may commission up to five deep dives: the part nobody can predict.
Static, dynamic, and the hybrid Tesserine runs
Static code decides
Dynamic the lead decides
Hybrid
Parallel or sequential follows from dependencies. Independent workers run in parallel, so the sweep takes the time of the slowest scout rather than the sum; Anthropic's Research system cut time by up to 90% on complex queries once its subagents, and their tool calls, ran in parallel. Dependent steps run in sequence: the deep dives need the sweep, and the citation check needs the draft. For fan-outs of dozens or hundreds of agents, Claude Code and the SDK also offer dynamic workflows. Despite the name, they are static orchestration: Claude writes a script once, and the runtime executes it.
Stop conditions and budgets belong at every level, enforced by the runtime rather than requested in a prompt. By default the Agent SDK allows three layers of nested subagents, 20 running at once and unlimited spend, so set all three, as the code below does. A maxTurns cap returns a worker's output marked partial, and the env caps bound depth and concurrency. At max_budget_usd the SDK refuses new subagents and ends the run with an error result.
from claude_agent_sdk import ClaudeAgentOptions, AgentDefinition
registry_scout = AgentDefinition(
description="Checks trial registries for one competitor. Use once per competitor.",
prompt=REGISTRY_SCOUT_PROMPT, # role, output schema, failure rule
tools=["mcp__registry__search", "mcp__registry__get_record"], # read-only
model="sonnet", # workers on a cheaper tier than the lead
maxTurns=15, # per-worker stop: output marked partial
)
options = ClaudeAgentOptions(
system_prompt=LEAD_PROMPT, # plan, delegate, merge, cite
mcp_servers=SOURCE_SERVERS, # registry, press and abstracts servers
allowed_tools=["Agent", "mcp__registry__*", "mcp__press__*", "mcp__abstracts__*"],
agents={"registry-scout": registry_scout, "press-scout": press_scout,
"abstract-scout": abstract_scout},
env={"CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH": "1", # workers cannot spawn workers
"CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS": "10"}, # at most 10 at once
max_budget_usd=40.0, # the whole run, subagents included
)
The last piece is failure. When one scout fails, say on a registry rate limit, never rerun all 45. Store each cell's result as it lands and retry only the failed cell, a limited number of times. If it still fails, ship the brief with the gap stated: "registry data for Orvane Bio unavailable this week". The lead must never fill that gap from memory. Claude Code and the SDK help here: a subagent that ends on an API error reports the failure rather than passing error text off as findings, a clean signal to record.
1.4.6 Seeing inside a multi-agent run
Two months in, Declan reports that the brief missed a new Orvane trial. Was it never briefed, missed by the scout, lost to a tool error, merged away as a duplicate, or cut by the citation check? One log line per run cannot tell you. Anthropic's Research team hit the same question and answered it with full production tracing.
Tracing starts with attribution. In the Agent SDK, every message from inside a subagent carries a parent_tool_use_id linking it to the delegation call that spawned it, so each tool call and message can be attributed to a worker. Ines records, per worker, the brief, the tools called, turns, tokens and a status (ok, empty, partial or failed), and per run the coverage of the 45 cells, duplicates merged, conflicts flagged and total cost.
Evaluation needs a different mindset. Two good runs can take different paths, one scout searching three times and another ten, so you judge the outcome and whether the process was reasonable, not whether it followed the steps you imagined. Start small: Anthropic's Research team began with about 20 queries drawn from real usage. For grading at scale, one model call acting as judge, scoring against a rubric from 0.0 to 1.0 with a pass or fail, matched human judgement best. Human review still catches what the judge misses: testers noticed early agents preferring content farms over authoritative sources.
Tesserine's eval set is 20 past weeks with the analysts' manual briefs, and Declan's missed trial joins it as a must-catch event. Here is the rubric; the first two lines can fail a brief outright.
Completeness: every must-catch event for the week appears. Score = found / expected; a missed phase 3 status change fails the brief.
Citation accuracy: every claim links to a source that states it; one unsupported claim fails the brief.
Factual accuracy: dates, phases and enrolment figures match the cited record.
Source quality: registries and company sources preferred over secondary coverage.
Efficiency: tokens and tool calls per competitor within budget; list every scout that hit its turn cap.
1.4.7 The exam traps
Most wrong answers here either add agents the requirement does not need or leave coordination to hope.
- ✗ Going multi-agent because the problem feels large or advanced. ✓ Split only for independent breadth, context worth isolating or distinct tools, and only where the task's value pays for several times the tokens.
- ✗ Splitting sequential phases of the same work (plan, draft, review one document) across agents. ✓ Keep work that shares context in one agent; each handoff loses some. A verifier that needs only the output and criteria is the split that works.
- ✗ One-line briefs such as "research Orvane Bio". ✓ State the objective, output format, tools, boundaries and effort. The brief is all the worker knows of the task.
- ✗ Several agents with no owner, no stop condition and no budget, trusted to sort out the work among themselves. ✓ An orchestrator owns routing and completion, and the runtime enforces turn, depth, concurrency and spend limits.
- ✗ Letting the lead decide a fan-out that is known in advance. ✓ Put required coverage in code, and spend model judgment where the path cannot be predicted.
- ✗ Rerunning everything when one worker fails, or letting the lead fill the gap. ✓ Keep each worker's result, retry the failed one a limited number of times, and report any gap that remains.
Four tempting fixes for a misbehaving fan-out
1.4.8 Put it together: design and break a small fan-out
You now have every piece of the decision. Build a three-worker fan-out, starve it first of its brief and then of its budget, and watch what breaks.
Decomposition techniques (1.5) go deeper on cutting a problem into pieces, which decides where your workers' boundaries fall. Business value pillars (1.6) turn the token multiplier into a cost-per-brief argument. Observability at scale (3.4) and monitoring (4.6) grow the per-worker trace into production dashboards, and connection protocols (3.7) take delegation across organisational boundaries.
Key takeaways
- ✓ Several agents pay off for independent breadth, context worth isolating and distinct tools; they cost several times the tokens of one agent and are harder to debug, so coupled work stays in one agent.
- ✓ Orchestrator-worker is the default pattern; pipelines, generator-verifier, handoffs and agent teams fit other structures of work, and production systems often combine them.
- ✓ A subagent sees none of the lead's context beyond its brief, so every brief states the objective, output format, tools and sources, boundaries and effort.
- ✓ Workers return compact, structured, cited results, and the lead merges them on a stable key, keeps every citation and surfaces conflicts.
- ✓ Code owns a fan-out that is known in advance; the lead owns what cannot be predicted; every level has a budget and stop condition the runtime enforces.
- ✓ A failed worker is one piece to retry or report, never a reason to rerun everything or to let the lead guess.
- ✓ Trace every worker to the delegation that spawned it, and evaluate runs on outcomes with a small rubric-based eval set and human review.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
33 CCAR-P questions on Domain 1, free
Every question in the bank is tagged to a domain, so you can drill 33 questions on Solution Design & Architecture alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 1 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.