Home › Study guides › CCAO-F › Domain 7 › Lesson 7.1
CCAO-F · Domain 7 · 10% of the exam · Lesson 7.1 · 19 min read
Diagnosing poor outputs: from symptom to root cause
How to trace a poor Claude output to its root cause: read the symptom, find where the problem enters, change one thing at a time and retest the same case.
Written against objective 7.1 of the official CCAO-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
7.1.1 Why the obvious fix rarely works
It's the monthly quality review. You are the knowledge manager at a regional broadband provider. For three months the support team has drafted help-centre answers in a shared Claude Project: an agent pastes in a customer's question, Claude drafts a reply, and the agent checks it and publishes. Kofi, the support team lead, brings three complaints. The answers are generic, and one even told a customer to "contact your internet provider", which is you. Some quote last year's prices. And most ignore the "maximum 120 words" rule, running to 250 words or more.
The quick fixes arrive before Kofi has finished. Switch to the most capable model. Put "IMPORTANT: FOLLOW ALL RULES" at the top of the instructions. Tell agents to ask for a shorter version every time. Each acts on what the output looks like, and none asks why it happened. Each also treats three complaints as one problem, when there may be three causes in three places.
This lesson is about troubleshooting: working back from a poor output to what produced it. The visible problem is the symptom. The root cause is the input or setting behind it, such as a missing fact, an old file or two rules that contradict each other. Fix the symptom and it returns in the next chat. Fix the cause and it stops.
Two ways to respond to a complaint
Treat the symptom
Find the cause
7.1.2 Read the symptom, then suspect the cause
Here is the belief that trips people up: a poor output means Claude isn't good enough at the task. Occasionally that's true. Far more often, Claude did what its inputs allowed, and the inputs were the problem.
Think of a damp patch on a ceiling. The stain is in the living room, but the leak is in the bathroom above, and repainting the ceiling buys you a few dry weeks. Outputs work the same way: the symptom shows up in the draft, and the cause sits upstream, in what Claude was given.
The good news is that symptoms are not random. Each points to a short list of likely causes, and Anthropic's Claude 101 course pairs several the same way. The table collects the pairs you'll meet most often.
| Symptom | Likely root cause | Usual fix |
|---|---|---|
| Generic or vague | Missing context about the situation or the audience | Say who it's for, what the product is and what the reader already knows |
| Wrong or outdated facts | A missing, outdated or unapproved source, or no instruction to answer from the approved one | Supply the current source and tell Claude to answer from it |
| Wrong format | Format never specified; no example | Describe the structure or show one example |
| An instruction ignored | Instructions that conflict across the prompt, Project and personal instructions; a rule buried or phrased only as a "don't" | Reconcile the conflict; state the rule plainly, with its reason |
| Inconsistent across runs | Criteria open to interpretation; no examples | Define the criteria and add a few varied examples |
| Forgets earlier details | A long, cluttered conversation or stale context | Start a fresh chat with a short summary of what matters |
| Too shallow for the task | The wrong model or feature for the job | Rule out missing context, then pick a more capable model or a fitting feature |
| Wrong numbers | Arithmetic done in the reply instead of by analysis; unclear data definitions | Define the metric and have Claude calculate it as code |
Memorise the pattern, not the wording: the symptom tells you which input to inspect first. Kofi's complaints give three different suspects: missing context, an outdated source, and conflicting or buried instructions. None of them is "the model isn't clever enough". A suspect is not a verdict, though. A generic answer can also come from an agent who pasted half a question, so the next two steps find where the problem really is.
7.1.3 Find where the problem enters
Every draft in the help-centre Project is built from several layers, and a problem can enter at any of them. Knowing them turns "the answers are bad" into a place you can open and read. A Project keeps its own instructions and uploaded knowledge for every chat inside it. Each person can also set personal instructions that apply to all of their own conversations. On Team and Enterprise plans, an owner can add organisation instructions for everyone, which take precedence over personal ones when the two conflict.
Five places a problem can enter
The quickest way to narrow it down is to ask where the symptom shows up. Every new chat in the Project starts from the same shared setup, so where a fault appears tells you where it lives.
| Where the symptom shows up | Where it probably lives | Look first at |
|---|---|---|
| One chat only | That chat | Its prompt and its history |
| Every chat, for everyone | The shared setup | Project instructions and knowledge |
| One person's chats only | What differs for that person | Their prompts and personal instructions |
| One topic only | The source for that topic | The file or data behind it |
The second row matters most. All three complaints appear across the whole team, week after week, which rules out one bad chat or one agent's prompt. Take the generic answers. The Project instructions say only "Draft help-centre answers for the support team." Nothing says that Claude writes for the provider itself, or who the customers are. So Claude writes as an outsider would, and sends customers to their provider. The fix goes where the gap is: a few lines of context in the Project instructions. A reminder typed into each chat, like asking for a shorter version every time, repairs one draft and leaves the cause in place for the next hundred.
7.1.4 Reproduce, inspect, change one thing, retest
Now diagnose the price complaint properly. The temptation is to jump straight to a fix. Resist it. A fix chosen before you have seen the fault is a guess, and a guess that seems to work teaches you nothing.
From symptom to fix in four steps
First, reproduce the fault. Take a question that produced an old price, such as "How much is Fibre 150 a month?", and ask it in a new chat in the Project. It comes back at £32, last year's price; this year's is £35. You now have a test case. Had the fault not appeared, the cause would more likely sit in the original chat than in the setup.
Second, inspect the inputs before touching any wording. Anthropic's prompting guide offers a useful test: would a colleague with minimal context be confused by what you gave Claude? Open the Project knowledge and the answer is plain: it holds two price lists, "Prices 2025" and "Price list from April", because nobody removed the old one. Claude can draw on both, and nothing tells it which is current.
Third, change one thing: remove last year's list and nothing else. Had you also rewritten the instructions and switched the model, you wouldn't know which change fixed the price. Finally, retest the same case in a fresh chat, so the only difference between before and after is your change; an older chat carries answers built on the old file. Then widen the check to a handful of real price questions, including an awkward one such as a discontinued plan.
7.1.5 Ask Claude what it understood, then check
The length complaint is harder: the 120-word rule is right there in the instructions, and Claude still breaks it. Here a diagnostic question helps. Asking Claude to restate its task, or to say why it answered as it did, can surface a misunderstanding that the draft alone never shows.
In the exchange below, look at Claude's reply: it quotes two rules and says how it chose between them. The last four lines are the fix you write once the clue is confirmed.
You: Before you draft anything, list the instructions you follow about answer length, and tell me what you do when they point in different directions.
Claude: Two instructions apply. Near the top: "Always list every troubleshooting step so the customer never needs to call back." Near the end: "Don't write long answers. Maximum 120 words." When a fix has many steps, I keep them all, because a customer who has to call back seems the worse outcome.
Rewritten rule, moved to the top of the Project instructions:
Keep every answer to 120 words or fewer, because customers read these answers on their phones while their connection is down.
If a fix needs more steps than fit, give the first three and link to the full guide.
When length and completeness conflict, length wins.
That reply is a strong lead, and it is only a lead. Claude's account of its own answer is written after the fact; it is not a recording of how the answer was produced. Anthropic's interpretability researchers found a clear case: asked how it added 36 and 59, Claude described carrying the one, while inside the model the sum was worked out quite differently. Treat an explanation like a witness statement: it tells you where to look, and you still check the scene.
So you check. The instructions hold the two lines Claude quoted, eleven lines apart, and the diagnosis has three parts. The rules conflict, and for a nine-step router fix they cannot both be met. The length rule is buried at the bottom of a long list. And it's phrased as a "don't", which says what to avoid but not what to do when a fix won't fit.
Anthropic's Help Center warns, in its advice on organisation instructions, that when instructions contradict each other, Claude may not follow either one reliably. So the rewrite settles the conflict with a tiebreak and, as Anthropic's prompting guide recommends, says what to do and why. Retest the same long router question in a fresh chat. If the draft lands under 120 words with a link, the diagnosis was right.
7.1.6 Wrong numbers and shallow answers
One more figure goes into Kofi's review: how readers rated last month's answers. You upload the ratings export, where each answer is marked helpful, not helpful or left blank, and ask Claude what share of answers were rated helpful. It replies 71%; the support dashboard says 83%. Rewording the question won't close that gap, and neither will a bigger model.
Two causes are worth checking, in this order. First, the data definition. The dashboard divides helpful ratings by rated answers only, while "share of answers" invites Claude to count the blanks too. Neither figure is wrong; they answer different questions, and only you or the dashboard's owner can say which one the business means. Second, the method. If Claude wrote the percentage straight into its reply without running any analysis, it did the arithmetic "in its head". As Anthropic's researchers put it, Claude was trained on text, not designed as a calculator, and a figure produced that way leaves no steps to check.
The same percentage, two ways
Written into the reply
Defined and calculated
The fix is to state the definition ("helpful ratings divided by rated answers; ignore blanks") and have Claude run the calculation with code execution. That feature lets Claude write and run code on an uploaded file, so each step is visible and repeatable. It is available on every plan, though a Team or Enterprise owner can switch it off. Then compare the result with a figure you trust, here the dashboard's 83%, before relying on new numbers.
Depth is the last suspect, not the first. If a draft is thin on a genuinely hard task, such as reconciling three overlapping outage policies, a more capable tier such as Opus may help. So may research mode, for a question that needs many sources. But rule out the inputs first. A more capable model handed an old price list quotes the old price more eloquently.
7.1.7 The exam traps
Troubleshooting questions describe a symptom, often one that keeps coming back, and offer several fixes. The wrong ones treat the symptom or change something beside the point.
- ✗ Switching to the most capable model when outputs go wrong. ✓ Diagnose first. No model can supply context nobody gave it, or tell a stale file from a current one without being told.
- ✗ Correcting a recurring problem chat by chat, or editing every draft by hand. ✓ Fix the Project's instructions or knowledge once and retest in a fresh chat; a problem across chats lives in the configuration.
- ✗ Repeating an ignored rule in capitals, or adding "IMPORTANT". ✓ Find the instruction it conflicts with, and state one rule with its reason and a tiebreak.
- ✗ Changing the prompt, instructions, files and model all at once. ✓ Change one thing and retest the same case, so you know what fixed it.
- ✗ Accepting Claude's explanation, or its rating of its own confidence, as the diagnosis. ✓ Use it as a clue, then confirm it by checking the inputs and retesting.
- ✗ Rewording the request when the numbers are wrong. ✓ Settle the data definition and run the calculation as analysis, checked against a known figure.
7.1.8 Put it together: trace a poor output to its cause
You now have the whole routine: read the symptom, find the layer, then reproduce, inspect, change one thing and retest, treating Claude's own explanation as a lead. The quickest way to make it instinct is to put a cause in on purpose and watch it come out.
The rest of Domain 7 builds on this routine. Adjusting your approach based on feedback and results (7.2) reads many results over time, so drift shows up before a quality review does. Optimising workflows for efficiency (7.3) then makes a setup that works faster and cheaper, which only pays once the causes of poor output are fixed.
Key takeaways
- ✓ A poor output is a symptom; the root cause is an input or setting upstream, and fixing only the symptom lets it return.
- ✓ Each common symptom points to a likely cause: generic to missing context, outdated facts to an old source, an ignored rule to conflicting or buried instructions.
- ✓ Find the layer where the problem enters (prompt, instructions, knowledge, feature or model, source data); a problem that recurs across chats lives in the configuration.
- ✓ Reproduce the fault, inspect the inputs first, change one thing, and retest the same case in a fresh chat.
- ✓ Claude's explanation of its own answer is a clue to check, not proof.
- ✓ Wrong numbers call for a clear data definition and a calculation run as code, not a reworded prompt or a bigger model.
Check your understanding
4 questions written for this lesson, then one from the CCAO-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAO-F questions on Domain 7, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Troubleshooting and Optimization alone, or sit the full 60-question timed simulator.
Open the CCAO-F question bank → Back to Domain 7 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.