Home › Study guides › CCAR-F › Domain 5 › Lesson 5.2
CCAR-F · Domain 5 · 15% of the exam · Lesson 5.2 · 21 min read
Escalation and ambiguity resolution
When a support agent should hand a case to a human, why sentiment and confidence scores mislead, and why it asks instead of guessing when a lookup is ambiguous.
Written against task statement 5.2 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
5.2.1 Why a capable agent still hands over the wrong cases
Picture someone in their first week on a shop's returns desk. A customer brings back an unworn jumper with the receipt, a case the staff handbook settles in one line, and the new hire calls the manager over. Ten minutes later another customer says a shop across town sells the same kettle for less. The new hire knocks the difference off on their own authority, although the handbook never mentions other shops' prices. Nobody would call this person unintelligent. They are guessing where the line runs because nobody drew it for them.
A support agent built on Claude meets the same problem in a sharper form. It has a tool, escalate_to_human, that passes the conversation to a person. It also has a system prompt, the standing instructions it reads before any customer writes, and that prompt often says little more than "escalate complex or sensitive cases". That is not a line; it is an invitation to guess. The guesses go wrong in both directions at once. The agent hands routine cases to people, which drags down first-contact resolution (the share of conversations it settles with no human involved). And it improvises on cases that genuinely needed a person's judgment.
A quieter kind of guess hides inside ordinary lookups. A customer gives their name, the database returns three accounts with that name, and the agent picks the one that "looks likely": the same mistake in different clothes. The cure for both is explicit escalation and clarification rules, written into the agent's instructions and shown in worked examples.
5.2.2 The three real reasons to escalate
Here is the question the rest of the lesson hangs on: what should actually make an agent call for a human? The instinctive answer is "the hard cases". It is wrong, and seeing why unlocks the whole task statement.
Let's set up the example we will follow to the end. An online homeware retailer runs a support agent built with the Claude Agent SDK (software development kit: Anthropic's ready-made library for building agents). Its four tools reach the retailer's customer and order systems through MCP (Model Context Protocol, an open standard for connecting AI applications to outside tools and data): get_customer, lookup_order, process_refund and escalate_to_human.
The target is 80% first-contact resolution; the agent sits at 61%. A week of transcripts shows two failures pulling in opposite directions. The agent escalated a customer who wanted the delivery fee back on a parcel that arrived six days late, which the refund policy covers in one rule. And it agreed, on its own authority, to match a competitor's lower price, although the policy covers only price drops on the retailer's own site.
Only three triggers hold up under scrutiny:
- The customer asks for a human. A clear request for a person is honoured, whatever the issue.
- A policy exception or a policy gap. The customer wants something the policy rules out (a return on day 41 of a 30-day window) or something the policy never mentions (matching a competitor's price). Either way, someone has to make a decision the agent was never authorised to make.
- No meaningful progress. The agent has tried the tools that apply and still cannot move the case forward: the order system keeps failing, or the fix needs an action no tool offers.
Notice what is missing: complexity. A billing dispute spanning three orders and two partial refunds is complicated. But if the policy covers every step and the tools can carry it out, the agent should handle it; that is what it is for. A calm one-line question can be the case that needs a person, because the policy is silent on it. Complexity measures how much work a case takes. The triggers ask whether the agent is ALLOWED and ABLE to do it.
| Situation | Trigger | What the agent does |
|---|---|---|
| Parcel six days late; policy refunds the delivery fee | None: covered, and a tool can do it | Resolves it with process_refund |
| Billing dispute across three orders, every step covered | None: complex but covered | Works through it and resolves it |
| "Can you match this competitor's price?" | Policy gap: policy covers own-site drops only | Escalates |
| Return asked for on day 41 of a 30-day window | Policy exception | Escalates; it does not grant exceptions |
lookup_order keeps failing; order status unknown |
No meaningful progress | Escalates with what it knows |
| "I'd like to speak to a person, please." | Explicit request | Escalates at once |
Memorise the three triggers, and notice that the busiest row in the table, the three-order dispute, is not one of them.
5.2.3 A request for a human is not the same as frustration
Two messages can sound alike and need opposite handling, and this is where many agents, and many exam candidates, slip.
The first: "I'd like to speak to a person, please." That is an explicit request, and the agent honours it immediately, without investigating first. It is tempting to be helpful ("let me just pull up your order"), but running lookups before handing over is refusing the request in slow motion. The customer has decided how this conversation should go, and that decision is theirs. The handoff passes on what the conversation already holds, such as the customer's name and what they asked for. What the agent must not do is start digging to see whether it could have solved the case itself.
Think of a hotel guest who asks the receptionist for the duty manager. A good receptionist picks up the phone rather than first trying to fix the complaint herself. The precise rule: an explicit request for a human triggers escalation before any investigation.
The second message: "Second late parcel this month. Honestly, this is ridiculous." That is frustration, not a request. If the issue sits within the agent's capability, the right move is to acknowledge the feeling in a sentence and offer the fix: "I'm sorry, two late deliveries in a month is frustrating. I can refund the delivery fee on this order right now; shall I go ahead?" Escalating here would swap an instant fix for a wait in a queue. But if the customer answers "No, just get me a person", they have now stated a clear preference. The agent escalates at once and does not try a second time to persuade them.
Frustration on its own never triggers a handoff, however strongly it is worded. The message to watch is the one in between, such as "Can someone actually sort this out?", which hints at a person without clearly asking for one. Treat it like frustration: acknowledge it and offer the fix. If the reply repeats the wish ("No, I mean a real person"), the preference is now clear, and the agent escalates at once, with no more lookups and no second pitch.
Three messages, three responses
Asks for a person
Frustrated, issue in scope
Calm, policy silent
5.2.4 Why mood and confidence scores are the wrong signals
Once a team accepts that its agent escalates badly, two shortcuts usually come up. Both produce a number, which feels objective, and both are unreliable stand-ins, or proxies, for what matters: whether this case needs a human.
The first is sentiment-based escalation: score each message for how negative it sounds and hand over below a threshold. Sentiment measures the customer's mood, and mood is not case complexity. In our transcripts, the furious customer with the late parcel was the easiest case of the week, and the polite customer asking for a price match was the one that needed a person. A sentiment router escalates the first and resolves the second, which is exactly backwards. Tracking sentiment across conversations is a reasonable way to judge how well the agent is doing; using it to decide who handles a case is the mistake.
The second is a self-reported confidence score: ask the agent to rate its own certainty from 0 to 1 and escalate when the number is low. The trouble is that the number is more generated text, produced by the same reasoning that made the decision, and nothing has checked it against real outcomes. In other words, it is uncalibrated. Our agent's failure was being sure where it should not have been. It granted the price match without hesitation, so it would have reported high confidence for exactly the case it got wrong. Ask the new hire from the returns desk "how sure are you?" and the one who gave away the discount will cheerfully say "very".
Cruder proxies fail for the same reason. A long message, or one typed in capitals, tells you about the customer's writing style, not about whether the policy covers their request.
Four tempting signals, one reliable rule
5.2.5 Several matches: ask, do not pick
Escalation is one half of this task statement. The other half is ambiguity: what the agent does when the data in front of it does not point to one answer.
A customer writes: "Hi, it's Sam Patel. I'd like to return the lamp I bought." The agent calls get_customer with the name, and three accounts come back. It is tempting to choose by a reasonable-sounding rule of thumb, a heuristic: the first in the list, the most recent order, the one with a lamp on it. Every one of those rules is a guess, and a wrong guess means a refund to the wrong person or another customer's order history shown to a stranger.
The right move is to ask for an additional identifier: something the customer knows and the system can match, such as an order number or the email on the account. The agent calls get_customer again with it and carries on only when exactly one account matches. It is also good practice not to read the candidates back ("are you the Sam Patel in Leeds or the one in Bristol?"). That question would show other customers' details to someone the agent has not yet identified.
Resolving several matches
get_customer returns 3 accountsThis is the escalation principle pointed at data instead of policy. In both cases the agent lacks something it needs to act correctly, and the honest response is to surface the gap rather than paper over it with an assumption.
5.2.6 The first fix: explicit criteria and examples in the prompt
The team now knows what is wrong. Which change comes first? Confidence thresholds and sentiment routers are already ruled out. A separate classifier model, trained on past tickets to predict escalations, might one day earn its place, but it needs labelled data, training and hosting. And the problem has a simpler cause: the prompt never told the agent where the lines are. A proportionate fix starts with the lightest change that addresses the root cause, and reaches for heavier machinery only if that change, once measured, falls short. Anthropic's advice on building agents says the same: find the simplest solution that works, and add complexity only when it is needed.
Here the lightest change is a rewrite of the system prompt, in two parts. The first is explicit escalation criteria: the three triggers, the things that are NOT triggers, and the rule for multiple matches, in plain words with the reason behind each. The second is few-shot examples, short worked cases that show the decision being made.
Anthropic's prompting guidance says a few well-crafted examples improve accuracy and consistency. It recommends three to five, relevant to the real task and diverse enough to cover edge cases. Each goes inside <example> tags, labels in angle brackets that mark where an example starts and ends, so Claude can tell examples apart from instructions. Diversity matters here. If every escalated example happens to feature an angry customer, the agent can learn the wrong pattern, that anger means escalate. So the most useful examples sit on the boundary: an angry customer with a covered request who is helped, next to a calm customer with an uncovered request who is escalated.
Here is part of such a prompt. The numbered list holds the criteria; the two <example> blocks show one case on each side of the line, each with its reason.
Call escalate_to_human when any of these is true:
1. The customer asks for a person. Escalate at once; look nothing up first.
2. The request needs a policy exception, or the policy does not cover it.
3. You have tried the relevant tools and cannot make progress.
Do not escalate because a case is long or the customer is upset. If the
policy covers it and your tools can do it, resolve it. If get_customer
returns more than one account, ask for the order number or account email;
never choose one yourself.
<examples>
<example>Customer: "Third late parcel. Refund the delivery fee, I'm fed up."
Covered by the late-delivery rule: acknowledge, then process_refund.</example>
<example>Customer: "Another shop sells this lamp for £15 less. Match it?"
Policy covers only drops in our own price, so it is silent: escalate.</example>
</examples>
In the Agent SDK this text goes into the prompt you pass as system_prompt in Python or systemPrompt in TypeScript. A support agent with its own name and persona normally uses a custom prompt string rather than Claude Code's built-in preset. Then measure: replay the week's customer messages and count correct escalations, unnecessary ones, and cases that should have reached a person but did not. Anthropic's customer support guide recommends tracking this kind of escalation accuracy.
5.2.7 The exam traps
Every trap in this task statement is the agent acting on the wrong signal: difficulty instead of policy, mood instead of a request, a guess instead of a fact. Here are the pairs.
- ✗ Escalating because a case is complex. ✓ Escalate on the three triggers. A complicated case the policy covers is the agent's job; a simple one the policy is silent on goes to a person.
- ✗ Investigating before honouring "I want a human". ✓ Escalate immediately. Lookups first, however well meant, override a decision that belongs to the customer.
- ✗ Escalating every frustrated customer. ✓ Acknowledge the frustration and offer the fix when the issue is in scope; escalate only if the customer then asks, or asks again, for a person.
- ✗ Routing on a sentiment score or a self-reported confidence score. ✓ Route on explicit criteria. Sentiment tracks mood and self-reported confidence is uncalibrated; neither tracks what the case needs.
- ✗ Picking the likeliest of several matching customers. ✓ Ask for another identifier and look up again. A wrong pick acts on someone else's account.
- ✗ Building a trained escalation classifier as the first fix. ✓ Add explicit criteria and few-shot examples to the system prompt first; it fixes the root cause at a fraction of the cost.
5.2.8 Put it together: calibrate a support agent's escalations
You now have every piece: the three triggers, the difference between a request and a feeling, the proxies to ignore, the clarifying question, and the prompt-level fix. The quickest way to make them stick is to run a small set of test conversations, watch the agent's decisions move as you change its prompt, and then break it on purpose.
The rest of Domain 5 builds on the same honesty. Error propagation across agents (5.3) is the multi-agent version of "cannot make progress". A subagent that fails has to say what failed and what it tried, so the coordinator can decide whether to retry, work around the gap or hand over. Human review workflows and confidence calibration (5.5) return to confidence scores and show what makes them trustworthy enough to route work on: calibration against labelled examples. Provenance in multi-source synthesis (5.6) applies ask-do-not-guess to sources: when two credible sources disagree, the system reports the disagreement instead of quietly picking a winner.
Key takeaways
- ✓ Escalate on three triggers: an explicit request for a human, a policy exception or gap, and no meaningful progress. Complexity alone is not a trigger.
- ✓ When policy is silent on a request, such as a competitor price match under a policy that covers only your own prices, the agent escalates instead of improvising.
- ✓ An explicit request for a human is honoured immediately, without investigating first.
- ✓ Frustration is not a request: acknowledge it, offer the fix when the issue is within the agent's capability, and escalate only if the customer then asks, or asks again, for a person.
- ✓ Sentiment scores and self-reported confidence are unreliable proxies for case complexity; route on explicit criteria.
- ✓ When a lookup returns several matches, ask for another identifier; never pick one by heuristic.
- ✓ The proportionate first fix is explicit escalation criteria plus few-shot examples in the system prompt.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
54 CCAR-F questions on Domain 5, free
Every question in the bank is tagged to a domain, so you can drill 54 questions on Context Management & Reliability alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 5 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.