Home › Study guides › CCAR-F › Domain 2 › Lesson 2.2
CCAR-F · Domain 2 · 18% of the exam · Lesson 2.2 · 22 min read
Structured error responses for MCP tools
Why a generic "Operation failed" wrecks an agent's judgement: the MCP isError flag, four error categories, isRetryable, and empty result versus access failure.
Written against task statement 2.2 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
2.2.1 Why "Operation failed" is the most expensive sentence a tool can say
Picture a new employee on a support desk. They ask a colleague to check an order, and the colleague comes back with "couldn't do it". Nothing else. Did the system crash? Was the order number typed wrong? Is the colleague not allowed to see it? The new employee cannot know, so they either ask again and again or give up and apologise to the customer. Either choice is a coin toss. The answer was not false; it was USELESS, because it left out the one thing that decides the next move: why.
Now meet a customer support agent built with the Claude Agent SDK, Anthropic's toolkit for building agents. It reaches your company's systems through four tools: get_customer, lookup_order, process_refund and escalate_to_human. The tools are served over the Model Context Protocol (MCP), an open standard for connecting AI applications such as Claude to tools and data.
One afternoon, four calls fail, each for a different reason. lookup_order times out because the order database is overloaded. Another lookup_order call is rejected because the order id ORD-4471x is malformed. process_refund refuses order ORD-4471 because it was delivered 45 days ago and the return window is 30. And get_customer is denied because the customer is also an employee, and the agent's credentials do not cover staff records. If every one of those comes back as Operation failed, the agent cannot tell them apart. It retries the refund refusal five times, and it abandons the timed-out lookup that would have worked on the second try.
The model is good at deciding what to do about a failure, as long as it knows what the failure IS. So the fix is not a cleverer prompt or a retry-everything wrapper. The fix is in the tool. It returns a structured error response: a result flagged as an error, carrying enough detail for the agent to reason about it.
The same afternoon, two kinds of tool
Generic errors four failures, one message
Structured errors four failures, four meanings
2.2.2 isError: how a tool tells the agent it failed
Start with the mechanism, because the exam names it. When an MCP tool finishes, it sends back a tool result: some content (usually one block of text) and an optional true-or-false flag called isError. Set isError: true and the agent knows the call did not succeed. Leave it out, and the content counts as a normal answer. That flag is the whole signalling convention.
An error result is still a result, not a crash. The agent loop does not stop. The result goes into the conversation like any other tool result (the Agent SDK adds it for you). On the next turn the model reads it and decides what to do: retry, try another tool, or explain the failure.
Why a flagged result rather than letting the code crash? Because the model never sees your program, only the conversation. Say the tool's handler (the function that runs when the tool is called) hits a problem and throws an exception, a program's built-in emergency signal. The Agent SDK can host your tools inside your own program, and that in-process MCP server catches the exception. It turns it into an error result carrying the raw exception text, so the loop survives. But a raw message like KeyError: 'order' tells the model nothing it can act on. Catching the failure yourself and returning isError: true with a message YOU wrote is how you control what the model reads.
The same flag exists one layer down with a different spelling. In the Messages API, where your own code runs the loop, a tool_result block carries is_error: true when the tool failed. The Agent SDK follows each language: Python's @tool decorator takes "is_error", TypeScript's tool() helper takes isError. One flag, two spellings, one meaning. The exam uses the MCP spelling, isError.
Here is lookup_order reporting a timeout the right way, in Python. Look at two lines: the except that catches the timeout, and is_error, which marks the result as a failure rather than odd-looking data.
@tool("lookup_order", "Look up an order by its id, e.g. ORD-4471", {"order_id": str})
async def lookup_order(args):
try:
order = await orders_db.get(args["order_id"], timeout=3)
except TimeoutError: # CATCH the failure yourself
return {
"content": [{"type": "text", "text":
"Order database timed out after 3 seconds. Retry after a short wait."}],
"is_error": True, # the FLAG: a failure, not data
}
return {"content": [{"type": "text", "text": json.dumps(order)}]} # success: no flag
That message already beats "failed": it says what happened and what to try next. Two sections from now, we turn that sentence into named fields the agent can rely on.
2.2.3 Four kinds of failure, four different next moves
Now the question that makes structure necessary: what should the agent DO with an error? There is no single answer, and that is the whole point. Failures fall into four categories, and each demands a different reaction.
- TRANSIENT. The system was briefly unable to answer: a timeout, a rate limit (a service refusing requests because too many arrived too fast), a service down for a minute. Nothing about the request was wrong, so the same call can succeed after a pause.
- VALIDATION. The input was wrong: a malformed id, a missing field, a date in the wrong format. The same call fails forever, but a CORRECTED call can succeed. The move is to fix the input, perhaps by asking the customer.
- BUSINESS. The input was fine and the system worked, but a rule said no: outside the return window, refund above a limit, account frozen. No retry or reformatting will change the rule. The agent's job is to explain it to the customer in words they can accept.
- PERMISSION. The caller is not allowed: the agent's credentials do not cover this account or operation. Retrying cannot change that. The move is to hand the case to someone who has access, which in our scenario means
escalate_to_human.
| Category | Example from the afternoon | Retryable? | What the agent should do |
|---|---|---|---|
| Transient | lookup_order times out |
Yes, after a pause, a limited number of times | Retry; if it keeps failing, report it as unresolved |
| Validation | lookup_order("ORD-4471x") is rejected |
Not as it stands | Correct the input (fix the format, or ask the customer) and call again |
| Business | process_refund refused, order outside the 30-day window |
No | Explain the rule to the customer and offer what is possible; never retry |
| Permission | get_customer denied for a staff record |
No | Escalate to a human; never hammer the tool |
Memorise the four names and the retryable column; the examples only need recognising. Notice that only ONE of the four is retryable as it stands. That is why "retry everything" is such a wasteful policy. In three categories out of four, it burns calls on something that can never succeed unchanged.
2.2.4 Structured metadata: category, isRetryable, and a message a person could read
If the category decides the move, the error must carry the category. That is what "structured" means here: not a paragraph of prose the model has to interpret, but a few named fields it can rely on. Three are worth knowing by name.
errorCategory. Which kind of failure it was:transient,validation,businessorpermission. The agent's choice of next move rests on this field.isRetryable. A boolean (a true-or-false value). It istrueonly when the same call could succeed later, which means transient failures. For everything else it isfalse, so the agent stops at once instead of learning by exhaustion. You will also see it writtenretriable: false; the meaning is identical.- A human-readable description. What went wrong and what to try next, in plain words. Anthropic's tool-use docs suggest "Rate limit exceeded. Retry after 60 seconds." instead of "failed". For a business error, add a second sentence the agent can pass STRAIGHT to the customer.
Be clear about where these fields come from. MCP defines the outer shape of a result: the content and the isError flag. It defines no error categories and no retry field. errorCategory, isRetryable and the customer message are fields YOU design and write inside the content. Usually they go in as a small block of JSON (JavaScript Object Notation, a standard way of writing named fields as text). These names are the ones the exam uses.
Generic error versus structured error
Generic what the tool returned
"Operation failed"Structured what the tool returned
errorCategory: businessisRetryable: falsestop at oncecustomerMessage"outside the 30-day return window"Here is what process_refund writes into the text of a result marked isError: true when order ORD-4471 is outside the window.
{
"errorCategory": "business",
"isRetryable": false,
"message": "Refund refused: order ORD-4471 was delivered 45 days ago; the return window is 30 days.",
"customerMessage": "This order was delivered 45 days ago, outside our 30-day return window, so I can't refund it automatically. I can raise a goodwill request for you."
}
With that in the conversation, the agent does not retry. It tells the customer why, offers the goodwill route or escalate_to_human, and moves on. Without it, the same agent retries process_refund until its turn limit stops it, then says "something went wrong". Same model, same prompt; the only difference is what the tool said.
2.2.5 Recover locally, propagate only what you cannot fix
The next question is WHERE recovery happens once agents are nested. In a multi-agent design, a coordinator agent splits a job and hands pieces to subagents: smaller agents, each with its own tools and its own conversation. Suppose the coordinator hands "gather everything about customer 88 and order ORD-4471" to a subagent holding get_customer and lookup_order. Then lookup_order times out. Who should retry? (The exam's research scenario asks the same thing about a search subagent whose web search gets rate-limited.)
The subagent. A transient failure is local knowledge: the subagent knows which call failed, with which inputs, and it has the tool in hand to try again. Bouncing the failure up to the coordinator means re-explaining the task, starting a new subagent and losing whatever the first one had collected. Anthropic's engineers make the same point about their multi-agent research system: restarts are expensive, so the system resumes from where the agent was when the error happened.
The subagent docs add the detail that makes local recovery cheap. Only a subagent's FINAL message returns to its parent; its intermediate tool calls and results stay in its own conversation. Three retries inside the subagent add nothing to the coordinator's conversation.
So the pattern is this. The subagent retries a transient error (isRetryable: true) a small, fixed number of times with exponential backoff, a pause that doubles each time: 1, 2, then 4 seconds. Anthropic's research team describes pairing the model's adaptability with deterministic safeguards such as retry logic. The model decides what to do about a failure; plain code enforces how many times. If a retry succeeds, the coordinator never hears of it. If the cap is reached, or the error was never retryable, the subagent stops and PROPAGATES: it reports the failure upward. "Never retryable" covers a validation error it cannot correct, a business rule and a permission denial.
What it propagates matters as much as when. A bare "lookup failed" recreates the generic-error problem one level up. The subagent's final message should carry three things. The partial results it did obtain: get_customer succeeded, here is the profile. The unresolved error with its category: lookup_order timed out on every attempt. And what was attempted: three retries, 1, 2 and 4 seconds apart. Now the coordinator can decide well: answer with what it has, try another route, or escalate.
Local recovery, then propagation
isError: trueisRetryabletrue only for transient2.2.6 "No orders found" is not an error
One last distinction, and the exam is fond of it because it is so easy to get wrong in code. A customer asks about order ORD-9999. The id is correctly formatted, and lookup_order searches the order database and finds no order with it. Is that a failure? No. The query ran, the database answered, and the answer is "no such order". That is a valid empty result: a successful call whose answer happens to be "nothing matches".
Compare it with the timeout earlier that afternoon, when the database never answered at all. That is an access failure: the agent has no idea whether the order exists. To a careless tool the two look identical, since in both cases no data came back. Reporting one as the other does damage in both directions. Report a valid empty result as an error, and the agent retries a query that already succeeded, then tells the customer "I couldn't check". Report an access failure as an empty result, and the agent confidently tells the customer their order does not exist, when in fact nobody could check. The second is worse: a wrong answer delivered with certainty.
The rule is honesty in the return value. A valid empty result has NO isError flag, and its text says "0 orders match ORD-9999", so the model sees the zero as a finding rather than a blank. An access failure has isError: true and, as always, a category and a retryable flag, because the agent now faces a retry decision. The question your tool must answer is "did I get an answer?", not "did I get data?".
Empty result or access failure?
isError; no answer: isError plus retry flag2.2.7 The exam traps
Every anti-pattern here is a tool that withholds information the agent needs, or an agent that ignores the information it was given. Questions describe the symptom (endless retries, a customer told something false, a coordinator restarting from scratch) and ask for the fix. The fix lives in the error response.
- ✗ Returning "Operation failed" for every error. ✓ Return
isError: truewitherrorCategory,isRetryableand a description. A uniform message removes the only information the agent needs to choose between retry, correct, explain and escalate. - ✗ Letting raw exceptions reach the model. ✓ Catch the failure and write the message yourself. The loop survives either way, but
KeyError: 'order'gives the model nothing to act on and carries no category or retry flag. - ✗ Retrying every error with exponential backoff. ✓ Retry only when
isRetryableis true. Backoff is right for transient errors and a waste for validation, business and permission failures, which cannot succeed unchanged. - ✗ Returning a policy refusal without a customer-friendly explanation. ✓ Add
retriable: falseand a customer message for business-rule violations. The agent's job is to tell the customer why, and it cannot invent the policy. - ✗ Sending every transient error up to the coordinator. ✓ Recover locally in the subagent; propagate only unresolved errors, with partial results and what was attempted. Restarting from the coordinator throws away work already done.
- ✗ Treating an empty result as a failure, or a failure as an empty result. ✓ Success with no matches gets no
isError; no answer getsisError: true. Mixing them either wastes retries or hands the customer a confident wrong answer.
2.2.8 Put it together: make the four failures legible
You now have every piece, from the isError flag to the honest empty result. The way to make it stick is to build a tiny tool, watch an agent misjudge a bad error, and then fix the error rather than the agent.
A good error message also works as a second tool description, because it teaches the model how to call the tool correctly next time. The rest of this domain builds outward from the tool. Distributing tools across agents (2.3) decides which subagent holds lookup_order, and so where local recovery can happen. MCP server configuration (2.4) is how those tools reach the agent at all. Later, error propagation across a whole multi-agent system (5.3) scales up the "recover locally, report honestly" rule, and escalation design (5.2) decides when a human takes over.
Key takeaways
- ✓ A uniform "Operation failed" hides the kind of failure, so the agent retries what cannot succeed and abandons what could.
- ✓ An MCP tool reports failure with a result flagged
isError: true(is_errorin a Messages APItool_result); the loop continues and the model reads the message you wrote. - ✓ Catch failures and write the error yourself; a raw exception forwarded by the SDK keeps the loop alive but gives the model nothing to act on.
- ✓ Four categories, four moves: transient (retry), validation (fix the input), business (explain to the customer), permission (escalate); only transient is retryable as it stands.
- ✓
errorCategory, anisRetryableorretriableboolean and a human-readable description are fields you design inside the error content, not part of MCP; business refusals add a customer-friendly explanation. - ✓ Retry transient errors locally inside the subagent, with a cap; propagate only unresolved errors, with partial results and what was attempted.
- ✓ A valid empty result is success without
isError; an access failure isisError: truewith a retry decision attached.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
66 CCAR-F questions on Domain 2, free
Every question in the bank is tagged to a domain, so you can drill 66 questions on Tool Design & MCP Integration alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 2 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.