Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 5 › Lesson 5.3

CCAR-F · Domain 5 · 15% of the exam · Lesson 5.3 · 21 min read

Error propagation across multi-agent systems

How a failing subagent reports to its coordinator: structured error context, access failure versus empty result, local recovery and coverage annotations.

Written against task statement 5.3 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.3.1 When one subagent fails, what does the coordinator know?

Picture a newspaper editor an hour before deadline, with three reporters out on one story. She can't see what they see; she knows only what each of them phones in. Then the city hall reporter calls with three words: "Couldn't get it." Was the office shut, or did a clerk say the figures don't exist? Did he get anything at all before he gave up, and is there another way in? She has an hour, several options and no idea which of them makes sense.

A multi-agent research system has exactly the same shape, and for a technical reason. A coordinator agent splits a research question into pieces and hands each piece to a subagent, a smaller agent that works in its own separate conversation. Everything the subagent does, from its searches to the pages it reads and the judgements it makes, happens out of the coordinator's sight. When it finishes, only its FINAL message goes back to the coordinator, as the result of the tool call that launched it, and every intermediate step stays behind, including every failed search.

So the failure report is part of the design, not an afterthought. Error propagation is how a failure travels from the agent where it happened to the agent that can decide what to do about it. Done well, the subagent fixes what it can on its own and reports the rest in a form the coordinator can act on. The coordinator then picks a recovery, and the final report tells its reader honestly where the evidence is thin.

5.3.2 Four ways to report the same timeout

Let's make it concrete with our research system, where a coordinator delegates to four subagents: web search, document analysis, synthesis (which combines the findings) and report (which writes the final document). A client asks how extreme summer heat is affecting hospitals in European cities, and the coordinator gives one web search subagent a single slice of that question: heat and hospital admissions since 2019. Its first search, for studies of heatwaves and admissions, returns six good sources, two of them national health ministry reports. Its second asks for official 2022 admissions statistics for Spain and Italy in one heavy request, and the search tool times out.

The subagent now has to tell the coordinator something, and there are four broad ways to do it. Three of them sound reasonable, which is exactly why an exam question may set all four side by side.

  • Generic status. The subagent retries a few times and then reports "search unavailable", which leaves the coordinator unable to tell which search failed, why, or whether the six sources are lost. Its only moves are blind ones: rerun the whole slice from scratch, or drop it.
  • Empty success. The subagent catches the timeout and reports the query as completed with no results. The coordinator reads "searched, nothing there" and sees no reason to recover, so the final report quietly lacks the official figures and nobody knows they were never checked.
  • Ending everything. Nobody handles the timeout, so it travels up to the top of the system, which stops the whole research run. Every slice the other subagents had already finished is thrown away over one slow query, when a second try or a different route could still have filled the gap.
  • Structured error context. The subagent reports the failure type, the query it attempted, the partial results it holds and the alternatives it sees, so the coordinator can choose a recovery with its eyes open.

Of the three wrong ones, the empty success is the most dangerous, because nothing looks wrong, and a confident report with a silent hole does more harm than one that admits a gap.

Four ways to report one timeout

"Search unavailable"failed, but what and why?
Empty result as successreads as "nothing exists"
Stop the whole rundiscards every finished slice
Structured error contexttype, query, partial results, alternatives
Three reports leave the coordinator guessing, misled or with nothing left; only structured context lets it choose a recovery.

5.3.3 Structured error context: what the coordinator needs to decide

What exactly belongs in the good report? Structured error context is a failure report with four named parts, and each part is there because it unlocks a decision the coordinator can't make without it. The failure type says what kind of failure it was, and what was attempted gives the exact query and how many tries were made. The partial results are whatever the subagent did manage to gather, and the alternatives are the routes it thinks could still work.

Think of the card a courier leaves when nobody answers the door. A good one says what happened (no one home), what was tried (rang twice), where your parcel is now (at the depot) and what you can do next (collect it or rebook). You can act on that card in ten seconds, while one that said only "delivery failed" would send you straight to the phone.

Anthropic's advice for a tool reporting an error to an agent makes the same point one level down: say what went wrong and what to try next, not a bare "failed". The table shows what each element carries in our timeout and which decision it unlocks.

Element In our timeout What it lets the coordinator decide
Failure type timeout: the search tool never answered It's an access failure, so a retry decision is on the table
What was attempted One query covering Spain and Italy for 2022, tried three times Don't repeat it word for word; a retry needs a changed query
Partial results Six sources, including two ministry reports Nothing to redo, and perhaps enough to proceed
Alternatives Split the query by country; read the figures from the ministry reports Try another route instead of the same one

Memorise the four elements; the examples only show why each one matters. Here is how the web search subagent's final message could record them: look at status, which separates a failed search from a completed one, and at the four fields on the failed query. The Claude Agent SDK (software development kit) that the system is built with doesn't fix these names; you define the shape in the subagent's prompt, as part of its output format.

{
  "slice": "Heat and hospital admissions in European cities, 2019 onwards",
  "queries": [
    {"query": "heatwave hospital admissions European cities studies",
     "status": "completed", "sources": 6},
    {"query": "official heat-related admissions statistics Spain Italy 2022",
     "status": "failed",
     "failure_type": "timeout",
     "attempted": "3 tries, waits of 2 s and 4 s, every one timed out",
     "partial_results": "the 6 sources above, 2 of them ministry reports",
     "alternatives": ["one query per country",
                      "document analysis reads 2022 tables in the ministry reports"]}
  ]
}

Now the coordinator, which holds the plan for the whole question, can choose between three recoveries. It can retry with a modified query: the failed search bundled two countries into one heavy request, so one query per country is a real change, not a repeat. It can try an alternative approach, asking the document analysis subagent to pull the 2022 tables out of the two ministry reports already found. Or it can proceed with the partial results if time or budget has run out. Anthropic's own research system works the same way, with a lead agent that reads what its subagents return and decides whether more research is needed.

5.3.4 No answer is not an answer of "none"

The subagent moves on to its third search, for studies of heatwave admissions in Reykjavik, and this one runs normally and finds no matching study. Now two searches have produced no sources, the timed-out one and this one, and although they look alike they mean OPPOSITE things.

The timeout is an access failure: the search never got an answer, so we've learned nothing about what exists, and the Spanish and Italian figures may be sitting on a government website right now. An access failure needs a retry decision: retry, reroute, or proceed and declare a gap. The Reykjavik search is a valid empty result: the query ran successfully and the answer is "none". That answer is information in its own right, since retrying won't change it and the report can state it as a finding.

Think of a blood test: if the clinic loses your sample, nobody tells you the result was negative; they ask you to come back. "Negative" is a result, while "sample lost" means there is no result yet, and the next step is completely different.

Confusing the two does damage in both directions. Report the timeout as an empty result and the coordinator accepts it as a finding, so the synthesis may even state that no official figures exist, a confident false claim. Report the Reykjavik search as a failure and the coordinator spends a retry, perhaps a whole new subagent, re-asking a question that was already answered. The report then flags a gap where there isn't one.

So the subagent labels every search's outcome explicitly, as either completed with zero matches or failed with a failure type. The search tool itself should already tell the two apart; the subagent's job is to carry that distinction up to the coordinator intact instead of flattening both into "nothing found".

Access failure versus valid empty result

Access failure the Spain and Italy query

The search never answered
Nothing learned about what exists
Needs a retry decisionretry, reroute or proceed
Unresolved: a gap in the report

Valid empty result the Reykjavik query

The search ran and answered
No match is the answerthat is a finding
No retry needed
In the report: "none found"
Both searches returned no sources, but one never got an answer and the other got the answer "none", so the coordinator treats them in opposite ways.

5.3.5 Recover locally, then pass up what is left

One question is left: who should retry the timeout? It's tempting to send every hiccup straight to the coordinator, since it's in charge, but resist that. The subagent knows the exact query, has the search tool in hand and already holds the six sources. The coordinator, by contrast, would have to brief a fresh subagent that starts with nothing but its instructions and has to redo the finished work. Anthropic's research team built their system to resume from where an agent was when an error hit, because restarts are expensive and frustrating for users.

So the subagent handles transient failures locally, meaning failures like this timeout that may clear on their own if you wait and try again. Your code wraps the search tool in retry logic that tries a small, fixed number of times and waits a little longer before each new try. Because only the subagent's final message returns, a retry that works never reaches the coordinator at all and adds nothing to its context. Our courier, after all, rings twice before writing the card.

What the subagent can't resolve, it propagates, together with what it attempted and its partial results. The retries must be bounded, because the decisions that remain need a view of the whole job that the subagent doesn't have. Is this slice important enough to deserve another subagent, does another slice already cover it, and is there budget left? Only the coordinator can weigh that, so the subagent owns the retry and the coordinator owns the trade-off.

Now look again at the generic status from earlier: its retries were exactly right. The mistake came at the very last step, when three attempts' worth of knowledge was squeezed into two words.

Local recovery, then propagation

Search times outthe Spain and Italy query
Retry in the subagenta few tries, growing waits
Still failingstop, keep the six sources
Report uptype, query, partial results, alternatives
Coordinator decidesretry, reroute or proceed
The subagent absorbs what a retry can fix; the coordinator hears only about what is still broken, with everything it needs to decide.

5.3.6 Coverage annotations: a report that shows its gaps

Follow the coordinator's choice: it retries with one query per country, Spain's figures come back, and Italy's search times out again just as the time budget runs out. So it proceeds with partial results, and now the danger moves downstream. The synthesis subagent writes fluent paragraphs, fluent paragraphs read as complete, and a reader has no way to tell that the Italian figures were never seen.

The fix is to make the output say how well each part is supported. Coverage annotations are labels in the synthesis output that tell the reader, topic by topic, which findings are well supported and which areas have gaps because a source was unavailable. The best old maps had a habit worth copying: where the surveyors never went, the map said so instead of sketching a likely coastline. A coverage annotation is that honest blank space.

The annotations don't appear by themselves. A synthesis subagent knows only what its brief contains, so the coordinator must pass along the gaps and the empty results as well as the findings. The synthesis subagent's output format then requires a coverage label for every topic area, and the report subagent carries those labels into the final document.

Topic area Coverage annotation What it rests on
Heatwaves and admissions, overall trend Well supported Six sources from several countries
Official admissions figures, Spain 2022 Supported, one source The national statistics release found by the split query
Official admissions figures, Italy 2022 Gap: source unavailable Search timed out on every attempt; figures not checked
Heatwave admissions in Reykjavik Searched, none found Search completed with no matching study

Memorise the two ends of the scale: well supported, and a gap because a source was unavailable. The Reykjavik row shows the earlier distinction surviving all the way to the reader, since a valid empty result appears as a finding, not a gap. Annotations are also what keep "proceed with partial results" honest. It's a legitimate choice only when the output admits it is partial; without annotations, it's the empty-success mistake moved one step later.

5.3.7 The exam traps

Every trap here either hides information the coordinator needs or throws away work it could have used. Questions usually describe the symptom, such as a report that is silently incomplete, a run that dies or a coordinator restarting from scratch, and ask for the fix.

  • ✗ Returning a generic "search unavailable" once the retries run out. ✓ Return the failure type, the attempted query, the partial results and the alternatives. The retries were fine; the two-word summary throws away what the coordinator needs to choose.
  • ✗ Catching the timeout and returning an empty result marked as success. ✓ Report it as a failure with structured context. An empty success reads as "nothing exists", so no recovery happens and the report is silently incomplete.
  • ✗ Ending the whole research run because one subagent failed. ✓ Let the coordinator decide per failure: retry with a changed query, try another route, or proceed. One slow query should not discard every slice that succeeded.
  • ✗ Reporting "no matches" and "no answer" the same way. ✓ Label a completed query with zero matches as a finding and a timeout as an access failure. Only the second needs a retry decision.
  • ✗ Sending every transient blip straight up to the coordinator. ✓ Retry locally with a bound, and propagate only what stays unresolved, with the attempts and the partial results.
  • ✗ Proceeding with partial results and writing the report as if it were complete. ✓ Add coverage annotations that mark well-supported findings and the gaps left by unavailable sources.

5.3.8 Put it together: follow one timeout from search to report

You now have the whole path. The subagent recovers locally and reports what it can't fix as structured context, keeping "no answer" apart from "none"; the coordinator chooses a recovery; the report marks its gaps. Building a tiny version and then breaking it on purpose makes the difference hard to forget.

The rest of this domain keeps returning to the same two instincts: don't lose what you know, and don't claim more than you know. Large codebase exploration (5.4) protects partial work in a different way, with state exports that let a crashed run resume instead of restarting. Human review and confidence calibration (5.5) decides which outputs a person should check before anyone relies on them. Provenance (5.6) is the natural partner of coverage annotations: where this lesson marks the areas the research couldn't reach, provenance keeps every claim tied to its source and marks where credible sources disagree.

Key takeaways

  • ✓ A coordinator sees only a subagent's final message, so any detail of a failure that the message leaves out is invisible to it.
  • ✓ A generic status hides what failed, an empty result marked as success hides that anything failed, and ending the whole workflow throws away what succeeded.
  • ✓ Structured error context names the failure type, what was attempted, the partial results and possible alternatives, so the coordinator can retry with a changed query, reroute or proceed.
  • ✓ An access failure (a timeout) needs a retry decision; a valid empty result (a query that worked and found nothing) is a finding and must be reported as one.
  • ✓ Subagents retry transient failures locally with a bound and propagate only what they cannot resolve, with their attempts and partial results.
  • ✓ When the system proceeds with partial results, the synthesis output carries coverage annotations separating well-supported findings from gaps left by unavailable sources.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

54 CCAR-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 54 questions on Context Management & Reliability alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources