Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 2 › Lesson 2.1

CCAR-F · Domain 2 · 18% of the exam · Lesson 2.1 · 22 min read

Tool descriptions that route correctly

Why thin tool descriptions make an agent pick the wrong tool, what a full one contains, when to rename or split a tool, and how prompt wording overrides it.

Written against task statement 2.1 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.1.1 Why the agent keeps picking the wrong tool

Picture a council office with two service counters. The sign above one says "Enquiries". The sign above the other also says "Enquiries". You have come about a parcel, so you pick a counter at random, explain yourself, and get sent to the other one. Nobody behind the counters did anything wrong. The signs gave you nothing to choose on, so you guessed.

Now the support agent. It handles returns, billing disputes and account issues. It reaches the company's systems through four custom tools served over MCP (the Model Context Protocol, a standard way of packaging tools for an agent): get_customer, lookup_order, process_refund and escalate_to_human. The first two were described in a hurry. get_customer says "Fetches a customer." and lookup_order says "Fetches an order." Each takes one text parameter called id. A customer writes "Can you check on 58213? Still nothing has arrived." Some of the time the agent calls get_customer with 58213, gets "no such customer" back, and then escalates or asks for details the customer already gave.

A human colleague would glance at the number, remember that five-digit references belong to orders, and open the order screen. Claude cannot. It has no view of the code or the database behind a tool, and no memory of earlier conversations, because each request starts fresh. All it has is text: each tool's name, description and input schema (the list of parameters the tool accepts), plus your system prompt (the standing instructions you give the agent) and the conversation. Of those, the description is the part that says what a tool is FOR, which makes it the model's main guide when it picks one.

2.1.2 What Claude sees when it chooses

Start with a question that sounds trivial: when the agent chooses between get_customer and lookup_order, how does that text actually reach it? When your application sends a request with a list of tools, the API builds a special system prompt from your tool definitions and places your own system prompt beside them. Claude reads that, plus the conversation so far, and decides. Tools on an MCP server arrive the same way: the server advertises each one as a name, a description and an input schema. The docs are explicit that Claude never sees your implementation, only the tool definitions and the results you send back.

The docs are just as plain about how the choice is made: Claude decides when to call a tool based on the user's request and the tool's description. Now look at what "Fetches an order." gives the model, in the left column below: a topic and nothing else. With two labels that vague, the model has nothing solid to choose on, so it leans on whatever weak cues the message happens to offer. That is why misrouting looks random in the logs. It is a guess, and guesses vary.

A thin description versus a full one

"Fetches an order."

A topic: ordersand nothing else
Reference formatunknown
What comes backunknown
Where get_customer beginsunknown

The full description

Returnsstatus, items, delivery, refunds
Inputfive digits, often written "#58213"
Example requests"my parcel", "my delivery"
Boundaryaccount questions go to get_customer
A one-line description gives the model a topic and nothing to choose on. A full one answers the questions the model would otherwise guess at.

When misrouting first shows up, the tempting fixes all work around the descriptions rather than on them. Few-shot examples (sample messages paired with the right tool) add tokens, the units of text the API processes and bills, to every request. They also leave the vague descriptions in place, so the next phrasing you did not list is still a guess. A routing layer in code that picks the tool by keywords takes the choice away from the model, and with it the language understanding that handles requests nobody anticipated. Merging the two tools into one is a legitimate redesign, but it is a rebuild, far more than a first step needs when the real problem is two one-line descriptions.

First responses to misrouting

Few-shot examplesextra tokens, root cause untouched
A routing layer in codebypasses the model's understanding
Merge into one toola rebuild, not a first step
A stronger prompt rulecan create new associations
Rewrite the descriptionsinputs, examples, edge cases, boundaries
The instinctive fixes all work around the model's decision. The right first step improves the text the decision is made from.

2.1.3 What a full description contains

So what goes into a description that routes reliably? Anthropic's tool-writing guidance offers a good test: describe the tool the way you would to a new hire on your team. Make explicit the context you would otherwise take for granted, such as how references are formatted, what a niche term means, and how one record relates to another. Four elements do most of the routing work: input formats, example queries, edge cases, and boundary explanations. Around them sits the frame: what the tool does and what it returns.

Element What it tells the model The failure it prevents
Purpose and output What the tool does and exactly what comes back Calling a tool for data it does not return
Input formats The shape of each parameter, and how people actually write it Sending an order reference to the customer tool
Example queries Requests, in customers' own words, that should trigger it Missing a request phrased in an unexpected way
Edge cases What happens when nothing matches or the input is unclear Retry loops and confident wrong answers
Boundaries When to use this tool and when to use the similar one Two tools competing for the same request

Memorise the four routing elements (input formats, example queries, edge cases, boundaries); they are the ones the task statement names. Recognise purpose and output as the frame around them. The docs are blunt about length: detailed descriptions are by far the most important factor in tool performance, so aim for at least three or four sentences per tool, more if the tool is complex. A long description is not a smell. A one-liner is.

Here is lookup_order rewritten. Three sentences carry the routing: "Use it when" holds the example queries, "Input" holds the format, and "It returns nothing about the customer's account" draws the boundary.

{
  "name": "lookup_order",
  "description": "Returns the status, items, delivery date and refund history of a single order. Use it when the customer asks about a specific purchase: 'where is my parcel', 'my delivery is late', 'the item arrived broken', or when they quote an order reference. Input: the five-digit order reference, such as 58213. Customers often write it as '#58213' or 'order 58213'; pass only the digits. It returns nothing about the customer's account; for name, email, plan or billing details use get_customer. If no order matches, it returns a not-found result: ask the customer to check the reference instead of retrying or passing it to get_customer.",
  "input_schema": {
    "type": "object",
    "properties": {
      "order_id": {"type": "string", "description": "Five-digit order reference, digits only, e.g. 58213"}
    },
    "required": ["order_id"]
  }
}

get_customer gets the mirror image: "Returns the account record of one customer: name, email, plan, billing status and open tickets. Use it when the request is about the person or their account, such as 'update my email' or 'why was I charged twice'. Input: a customer id such as C-40017, or the email address on the account. A bare five-digit number is an order reference, not a customer id; use lookup_order for anything about a specific order." Now run the parcel message past both. "Nothing has arrived" matches lookup_order's example requests, 58213 matches its input format, and get_customer says outright that a five-digit number is not for it. Three independent signals point the same way, and the guess is gone.

Two smaller points. Parameter names and their descriptions are read too, so rename a vague id to order_id, as Anthropic's tool-writing guidance advises, and describe its format. And do not confuse the example queries inside a description with the optional input_examples field of a tool definition. That field holds sample inputs, such as {"order_id": "58213"}, that must match the schema. The docs suggest it for complex or format-sensitive inputs and say to prioritise the description.

2.1.4 Removing overlap: rename and split

Sometimes the descriptions are not thin so much as twins. The clearest case comes from a multi-agent research system, where a coordinator agent hands work to four specialist helpers, called subagents: search, analysis, synthesis and report. Suppose the analysis subagent has two tools, analyze_content and analyze_document. Their descriptions read "Analyzes content and returns key findings" and "Analyzes a document and returns key findings". One was built for web search results, the other for uploaded files, but neither the names nor the descriptions say so. Web results land in the document tool, PDFs land in the content tool, and the findings come back garbled because each tool is fed the other's input.

The first repair is to rename, so the name itself carries the difference. analyze_content becomes extract_web_results. Its new description says it takes the results of a web search and returns the relevant passages with their titles and URLs, and that uploaded files belong elsewhere. The name is part of the text the model reads, so "web" in the name and "web search results" in the description now agree. And "analyze" no longer sits in both names, inviting a coin toss. Anthropic's write-up of its own research system makes the same point: bad tool descriptions can send agents down completely wrong paths, so each tool needs a distinct purpose and a clear description.

The second repair is for one tool doing several jobs. Look closer at analyze_document: it quietly handles three different requests. Pull the figures out of a report. Summarise it. Check whether a claim is supported by it. One vague description covers all three, so the model has to guess which mode you meant, and the output shape changes with the guess. Nothing downstream can rely on it.

Split it into purpose-specific tools, each with a defined input/output contract: a fixed statement of what goes in and what comes out. The first, extract_data_points, takes a document and the fields wanted, and returns values with their locations. The second, summarize_content, takes a document and a target length, and returns prose. The third, verify_claim_against_source, takes a claim and a document, and returns supported or not, with the passage. One purpose, one input shape, one output shape per tool.

Splitting one generic tool into three purpose-specific tools

analyze_document

Any document, any request
Mode inferred from wording
Output shape varies

extract_data_points

In: document + fields wanted
Out: values with locations

summarize_content

In: document + target length
Out: prose summary

verify_claim_against_source

In: claim + document
Out: supported or not, with evidence
The generic tool forces the model to guess a mode and gives a different output shape each time. Each replacement has one job and one contract.

Back to the support agent, because the same diagnosis applies. Are get_customer and lookup_order twins? No: they have distinct purposes and distinct outputs, so they stay two tools and only their descriptions needed work. Is either one secretly three tools? No: each returns one kind of record. So the diagnosis has three outcomes. Distinct tools with thin descriptions: describe them. Overlapping names and descriptions: rename and re-describe. One tool whose modes take different inputs and return different outputs: split it, with a contract for each piece.

2.1.5 When the system prompt overrides a good description

Now the failure that catches teams who did everything above. You rewrote the support agent's descriptions, and misrouting fell sharply. Yet two stubborn slices remain. Customers who write "My parcel 58213 is a week late" are handed straight to escalate_to_human, although lookup_order could answer them in one call. And "My account says order 58213 was delivered, but nothing came" goes to get_customer, although it is about an order. The descriptions are fine. The system prompt contains two lines written months earlier: "Delivery problems are our top priority: escalate them quickly." and "Whenever a customer mentions their account, use get_customer."

Read those lines with the model's eyes. Whoever wrote the first one meant "deal with these fast". But "escalate" is, almost word for word, the name of a tool, so the line also reads as "send delivery problems to escalate_to_human". The second line is a keyword rule: it fires on the word "account", not on what the customer wants. The exam's term for both is a keyword-sensitive instruction: wording in the system prompt that ties a word to a tool and so creates an unintended tool association. The system prompt sits in the same assembled prompt as the tool definitions, and the docs confirm that prompting influences which tool Claude picks.

Picture an office where "escalate" has one meaning: hand the case to a supervisor. A manager says "delivery complaints are urgent, escalate them quickly", meaning "move fast". Staff hear an instruction about the supervisor, and every complaint lands on the supervisor's desk. The precise point: an instruction in the system prompt can outweigh a well-written tool description, so misrouting that survives a description rewrite usually lives in the prompt.

The fix is a review, and it is a skill the exam tests. Read the system prompt line by line and flag three kinds of wording:

  • A tool's name or description, used as an ordinary word, like "escalate".
  • Keyword rules of the form "when the customer mentions X, use Y".
  • "Always" rules that point at one tool whatever the request.

Then rewrite each flagged line around the customer's intent and the outcome you want, and let the descriptions do the routing.

Keyword-sensitive line The association it creates Rewrite around intent
"Delivery problems are our top priority: escalate them quickly." "Escalate" is a tool name, so delivery complaints go to escalate_to_human "Resolve delivery problems quickly. Hand a case to a person only when policy does not cover it or the customer asks for one."
"Whenever a customer mentions their account, use get_customer." Fires on the word "account", even in order questions "Look up the customer when the request is about the person or their account details."
"Anything with a number in it is an order question." Fires on digits: "I was charged twice on account C-40017" goes to lookup_order Delete it: lookup_order's description already says which references it takes

Memorise the shape of the fix, not these exact lines: find wording that points at a tool or fires on a word, and restate it as intent.

2.1.6 The exam traps

Every anti-pattern in this task statement either leaves the model to guess or patches the guess somewhere other than the text the model chooses from. Questions present a misrouting symptom and ask for the fix.

  • ✗ One-line descriptions on similar tools. ✓ Expand each with input formats, example queries, edge cases and boundaries. The description is the primary mechanism for tool selection.
  • ✗ Reaching first for few-shot examples, a keyword router or a merged tool. ✓ Rewrite the descriptions first. Examples add tokens and leave the root cause in place, a router bypasses the model's understanding, and merging is a rebuild a first step does not need.
  • ✗ Describing what the tool does and stopping. ✓ Say what it does NOT do and which tool to use instead. The boundary sentence is what separates two similar tools.
  • ✗ Two tools with generic names and near-identical descriptions. ✓ Rename so the name states the domain and re-describe with domain-specific detail: analyze_content becomes extract_web_results.
  • ✗ One generic tool whose output depends on how the request is phrased. ✓ Split it into purpose-specific tools with defined input/output contracts: extract_data_points, summarize_content, verify_claim_against_source.
  • ✗ A system prompt that uses a tool's name as an ordinary word, or ties a tool to a keyword. ✓ Review it for keyword-sensitive instructions and rewrite them around intent, because an instruction can override even a good description.

2.1.7 Put it together: describe, rename, split, then read the prompt

You now have every piece: what the model reads, what a full description contains, when to rename or split, and how one prompt line can undo it all. The quickest way to make it stick is to cause the misrouting yourself, fix it, and then break it again from the prompt side.

The rest of Domain 2 builds on descriptions that route. Structured error responses (2.2) are what lookup_order should return when 58213 matches nothing, so the agent can act instead of retrying. Distributing tools across agents and configuring tool_choice (2.3) decides which of these tools each agent sees at all, and whether a tool call is required. MCP integration (2.4) is where these descriptions live: in the tool list an MCP server offers to Claude Code and the Agent SDK. The built-in tools (2.5) arrive with descriptions already written; your job there is to choose between them.

Key takeaways

  • ✓ Claude picks a tool from text alone (each tool's name, description and input schema, plus your system prompt and the conversation), and the description is the primary mechanism for that choice.
  • ✓ Minimal descriptions on similar tools turn selection into a guess; the most effective first step is to expand them, not to add examples, a routing layer or a merged tool.
  • ✓ A full description states what the tool does and returns, its input formats as people actually write them, example requests, edge cases and its boundary with similar tools, in at least three or four sentences.
  • ✓ Overlapping names and near-identical descriptions cause misrouting; rename so the name states the domain (analyze_content to extract_web_results) and re-describe with domain-specific detail.
  • ✓ A generic tool whose modes take different inputs and return different outputs should be split into purpose-specific tools with defined input/output contracts.
  • ✓ Keyword-sensitive wording in the system prompt can override a good description; when misrouting survives a rewrite, review the prompt and rewrite such lines around intent.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

66 CCAR-F questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 66 questions on Tool Design & MCP Integration alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources