Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 4 › Lesson 4.4

CCAR-P · Domain 4 · 16% of the exam · Lesson 4.4 · 23 min read

Diagnosing system issues: prompt failures, hallucinations and model mismatch

Trace a bad answer to the layer that caused it: reproduce, isolate, compare, then tell prompt failures, hallucinations, lookalikes and model mismatch apart.

Written against objective 4.4 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

4.4.1 Why "the model got worse" is not a diagnosis

Coldharbour sells a contract-management platform to large companies, and every sizeable deal brings a security questionnaire: 150 to 300 questions on encryption, access control, certifications and incident response. An assistant now drafts each answer from Coldharbour's knowledge base: security policies, a certifications register and past answers the security team approved. It searches the base through a search_kb tool: Claude requests a search, and the application runs it and returns the matching passages. Thorsten, who leads security and compliance, signs off every draft before a customer sees it.

In one week, sales reports three problems. A draft told a public-sector customer that Coldharbour holds FedRAMP authorization, the US government's security authorization for cloud services. It has never had one, and Thorsten's reviewer caught the claim a day before sending. Procurement portals reject answers because drafts ignore the rule that every answer opens with Yes, No or N/A. And since the assistant moved from Claude Sonnet 4.6 to Claude Sonnet 5.5 two weeks ago, the sales engineers say the answers feel weaker. The head of sales has a one-line fix: move everything to Opus, a more capable and more expensive tier.

Mirela, Coldharbour's solution architect, will not decide that on a hunch, because three complaints in one week say nothing about whether they share a cause. The model sits in the middle of a pipeline: the incoming question, the search, the prompt, the model and its settings, the tools, and the code that turns its reply into a draft. A bad answer can start in any of those layers, and a bigger model changes only one. Diagnosis means finding the layer that failed before you change anything.

One fix for three complaints, or one diagnosis each

The tempting fix

Three complaints in one week
"The model got worse"
Move everything to Opushigher cost, causes untouched

The diagnosis

Reproduce each complaintwith its exact inputs
Trace it to a layerprompt, grounding, settings
One targeted change eachproven on the same cases
A bigger model treats three different faults as one; tracing each complaint to its layer finds three smaller, targeted fixes.

4.4.2 Reproduce, isolate, compare, change one thing

Here is the habit that separates diagnosis from guessing: change nothing until you can make the failure happen again. "The answers feel weaker" cannot be fixed, but a specific question that produced a specific wrong draft can. So Mirela asks sales for the request IDs of the failing drafts, not for more opinions.

Each ID leads to a trace, the stored record of that one request. It holds the replay inputs: the question, the passages search_kb returned, the prompt version, the model ID and settings, and every tool call and result. It also records the stop_reason, the field in every API response that says why the model stopped, and whether a cache answered. A good mechanic works the same way: make the car make the noise again, then disconnect one component at a time until it stops.

  1. REPRODUCE. Rerun the request from its recorded inputs a few times, because outputs vary from run to run. If it never fails again, suspect something you did not record, such as a cache or a transient tool error. When the trace already shows the failing step, reading it counts as reproducing.
  2. ISOLATE. Swap one layer for a known good version and hold the others fixed.
  3. COMPARE. Diff the failing configuration against the last known good one: which prompt, model, settings or knowledge-base version changed, and when.
  4. CHANGE ONE THING. Rerun the failing cases plus a sample that used to pass, so you learn what fixed it and what it broke.
Layer Isolation move If the draft is now right
Input Rerun with the question exactly as the customer sent it Your input handling altered the question
Retrieval Hand the model the passage that holds the answer, skipping search Search missed or misranked it
Prompt Run the previous prompt version on the same inputs The prompt change caused it
Model Run the same frozen inputs on the previous model The model or its settings changed behaviour
Tools Rerun with known good tool results in place of the recorded ones A tool fed the model bad or missing data
Post-processing Read the raw API response instead of the draft the user saw Your code cut, cached or mangled the reply

Behind the table sit three questions: what changed, what did the model actually see, and what did it return? Two cheap checks already split Coldharbour's complaints. Replayed on Sonnet 4.6, the FedRAMP draft fails identically, so the switch did not cause it. The format failures began with a prompt edit in July, weeks before the switch. Only the "weaker" complaints start on switch day, so the three complaints point to three separate causes.

4.4.3 Prompt failures: when the instructions lose

The format rule is in the prompt, so why does the model ignore it? Usually because something else in the prompt says the opposite, more convincingly. A prompt failure is output that breaks a requirement because the instructions themselves are contradictory, hidden, incomplete or undermined, and each kind leaves a different symptom.

Failure What you see Fix
Conflicting instructions One rule obeyed, another broken, varying from case to case Remove the conflict, or say which rule wins and why
Buried instructions A rule followed on short inputs, dropped on long ones Give the rule its own labelled section, after long documents
Missing context Generic answers or guessed specifics Supply the facts, or a tool that fetches them
Ambiguous format Several shapes for the same kind of answer Define the format exactly, or enforce it with a schema
Misleading examples Output copies the examples, not the instruction Make the examples demonstrate the rule, varied and tagged

Mirela compares prompt versions and finds the edit that marketing made in July. Look at the first line, the rule nine paragraphs down, and how every example opens.

You are Coldharbour's security expert. Write thorough, narrative answers that showcase our security maturity. Answer every question fully and positively; incomplete answers cost us deals.
[... 1,800 words on tone, product names and escalation contacts ...]
Formatting: start each answer with Yes, No or N/A.
<example>Q: Do you encrypt customer data at rest? A: Coldharbour maintains a comprehensive encryption programme in which all customer data...</example>
<example>Q: Do you run annual penetration tests? A: Coldharbour partners with independent specialists who...</example>
<example>Q: Is production access reviewed? A: Coldharbour's access governance framework ensures...</example>

Three failures stack up. The narrative instruction conflicts with the format rule, which is buried among 2,000 words about tone and contacts. And all three examples open with "Coldharbour". Anthropic's prompting guide calls examples one of the most reliable ways to steer format, which cuts both ways: here they steer toward the wrong one.

She changes one thing at a time on 40 failing questions. First she deletes the narrative line and moves the format rule into its own section, with its reason: portals read the first word. Then she replaces the examples with three varied ones opening Yes, No and N/A, and each rerun cuts the failures. But the consumer here is a portal's parser, not a person reading prose. So she enforces the shape with structured outputs, a JSON schema the API constrains the reply to, shown in the next section because it also fixes the first complaint.

4.4.4 Hallucinations: find where the fact went missing

Why would a model invent a certification? A hallucination is a fluent, confident answer that nothing in the model's input supports. The model produces the most plausible answer to what it sees, and when the fact is absent and the prompt demands a complete, positive answer, "Yes, FedRAMP Moderate" is plausible. So the diagnostic question is not why the model lied, but where the fact went missing.

Where the missing fact went

Not in the knowledge base

Search found nothingcorrectly
An answerability gap
Allow "not held", require evidence

In the base, not in context

Search missed itor a tool failed
A retrieval or tool fault
Fix the search, not the prompt

In context, contradicted

The passage was there
A prompt or model fault
Quote first, then test a stronger tier
Two questions settle it: is the fact in the knowledge base, and was it in the context the model saw? Each answer points to a different layer and fix.

The FedRAMP trace settles it in minutes. search_kb returned the certifications register, which lists SOC 2 Type II and ISO/IEC 27001, and a past answer about SOC 2 scope. Nothing mentioned FedRAMP, because Coldharbour has none, so retrieval worked. The cause is the prompt: "answer every question fully and positively", with no permitted way to say "not held". Nor does the prompt say the register is complete, so the model cannot read FedRAMP's absence as a No. A bigger model would meet the same gap and the same pressure to fill it.

Anthropic's guide to reducing hallucinations lists four prompt-level tools. Give the model permission to say it does not know. For long documents, have it extract word-for-word quotes before it answers. Ask for a supporting quote behind each claim, and retract any claim without one. And restrict it to the documents provided rather than its general knowledge. These reduce hallucinations, the guide notes, without eliminating them.

Mirela applies them: the prompt now permits an honest answer, states that the register lists every certification Coldharbour holds, and asks for a quote behind each claim. She also adds a check that does not depend on the model: certification names are a closed set, and her code blocks any draft that names one missing from the register. Here is the schema she sends in output_config.format. Look at the answer enum, which includes an honest way out, and the evidence array, which her code checks quote by quote against the passages the search returned.

{
  "type": "json_schema",
  "schema": {
    "type": "object",
    "properties": {
      "answer": {"type": "string", "enum": ["Yes", "No", "N/A", "Unverified"]},
      "explanation": {"type": "string"},
      "evidence": {"type": "array", "items": {
        "type": "object",
        "properties": {"passage_id": {"type": "string"}, "quote": {"type": "string"}},
        "required": ["passage_id", "quote"],
        "additionalProperties": false
      }}
    },
    "required": ["answer", "explanation", "evidence"],
    "additionalProperties": false
  }
}

The API also has a citations feature that returns the exact passage behind each claim, but one request cannot combine it with structured outputs; the API returns a 400 error. The portal's parser needs the guaranteed shape, so Mirela keeps the schema and carries the evidence in its fields. Her code marks any draft whose quotes fail the check "Unverified" and sends it to Thorsten's team.

4.4.5 Lookalikes: truncation, swallowed tool errors and stale caches

Some failures look like the model misbehaving when it did the sensible thing with what it received, or never saw the request at all. Coldharbour's week held three.

Truncation. A draft for a 14-part data-retention question stops mid-list, as if the model gave up. The trace shows stop_reason: "max_tokens": the reply hit max_tokens, the output cap the request sets, and the application showed the fragment as finished. Your code must read stop_reason before using a reply; for a final draft, only end_turn means the model finished. On max_tokens, raise the limit or continue the response, and never ship the fragment. With a schema, a truncated reply may not even match it.

A tool error swallowed as an empty result. A draft says Coldharbour has no documented incident-response plan, which it has. It looks like a hallucination, but search_kb had timed out, and the wrapper caught the exception and returned an empty list. The model read "nothing found" and reasoned correctly to a wrong conclusion. Return the failure as a tool_result, the block that carries a tool's output back to the model, with is_error: true and a message saying what happened. The model can then say it could not check, and you can count errors apart from real empty results.

A stale cache. A draft quotes the 90-day password-rotation rule that a new policy replaced in August. The trace shows a cache hit in 40 milliseconds, word for word a July draft, because Coldharbour's own response cache keys on the question text alone. Put the knowledge-base version in the key, or clear entries when a policy changes. Anthropic's prompt caching is not the suspect: it never stores answers, only the processing of an identical prompt prefix, so an edited policy misses the cache.

Symptom Layer to check first Evidence that settles it
A confident claim nothing supports Prompt and grounding The search found nothing relevant, and the base holds no such fact
Wrong answer though the base holds the right passage Retrieval The passage is absent from the results, or an old version came back
Right passage in context, answer contradicts it Prompt, then model Fails on a clean prompt; clears on a stronger tier
A rule or format ignored Prompt Began with a prompt version; a conflicting rule or example in it
Answer stops mid-sentence or mid-list Output settings stop_reason is max_tokens
"We have no..." where the record exists Tools The tool log shows a timeout returned as an empty result
Outdated content, identical to an old answer Your own cache A cache hit, near-zero latency, a key with no data version
Quality drop from the day of a model change Model and its settings Old and new models differ on the same frozen inputs

Memorise the third column: each layer leaves a different fingerprint in the trace.

4.4.6 Model mismatch: prove it with a replay, then decide

"Answers got weaker when we changed models" sounds like a verdict on the new model, but it is a hypothesis. Model mismatch takes two forms: a task beyond the tier you chose, or behaviour that changed with the new model through prompts tuned for the old one, moved defaults or removed features. Both are tested with a replay: the same frozen inputs run through the old and new model and scored against approved answers, so only the model varies.

Mirela replays 120 questions from last quarter with frozen search results. Sonnet 5.5 wins seven of nine categories and loses in two places. First, 11 drafts end with max_tokens, the data-retention draft among them: the truncation lookalike, triggered by the switch. The 1,200-token cap was sized for Sonnet 4.6, which ran without thinking unless asked. Sonnet 5.5 thinks by default, reasoning in thinking blocks before it answers, and that thinking counts toward max_tokens. Its tokenizer also turns the same text into about 30% more tokens. The migration guide says it plainly: revisit max_tokens.

Second, cryptography answers lost detail. The prompt says "for encryption questions, name the algorithm and key length"; Sonnet 4.6 stretched that to TLS and hashing, and Sonnet 5.5 does what it is told. Anthropic's prompting notes for Claude Sonnet 5 describe this, and the Sonnet 5.5 guide calls them a reasonable starting point. Those notes say the model reads prompts literally and does not silently generalise an instruction from one item to another. The fix is to state the scope. Effort, the setting that trades depth of thinking against speed and cost, is recalibrated too, so she reruns the replay at medium and high effort rather than trusting the default.

Kind of mismatch Its signature Fix
Task beyond the tier Errors cluster in the hardest questions, survive prompt fixes, clear on a stronger model with the same inputs Route that slice to a stronger tier, justified by the eval and cost per answer
Prompt tuned for the old model Starts on switch day; traces to specific instructions Rewrite them per the new model's prompting guide, then replay
Changed defaults and limits Silent: truncation, new response shapes, shifted cost Apply the migration guide: max_tokens, effort, reading blocks by type
Removed features Loud: requests fail with 400 errors Replace them before switching, as the migration guide lists

The loud row is the easy one. On Sonnet 5.5, five settings return a 400 error: a thinking budget, non-default sampling values (temperature, top_p, top_k), a prefilled assistant reply, a forced tool_choice, and thinking switched off with "disabled". A capability gap would look different. The losses would sit in the hardest category, such as multi-region data-residency questions, survive these fixes, and clear on a stronger tier with the same frozen inputs; Mirela would then route that category up a tier and price it per answer. Here, after three changes, Sonnet 5.5 matches or beats Sonnet 4.6 everywhere, and she records the decision.

Decision: keep Claude Sonnet 5.5 for questionnaire drafts.
Evidence: replay of 120 approved answers with frozen inputs; after three changes, Sonnet 5.5 matches or beats Sonnet 4.6 in all nine categories.
Changes: max_tokens raised from 1,200 to 4,000; the algorithm rule scoped to encryption, TLS and hashing; effort kept at high after sweeping medium and high.
Rejected: rollback to Sonnet 4.6 (gives up gains in seven categories); Opus for everything (higher cost, and neither cause was capability).
Revisit: if any category falls below its target on the monthly replay.

4.4.7 The exam traps

Each trap below acts before anyone knows which layer failed, and each has the same cure.

  • ✗ Moving to a bigger model because answers got worse. ✓ Reproduce and isolate first: a bigger model gets the same missing passage, conflicting prompt and max_tokens cap.
  • ✗ Adding "be accurate, never invent facts" to stop hallucinations. ✓ Find where the fact went missing, allow an honest "not held", require evidence, and check closed-set claims in code.
  • ✗ Changing the prompt, the model and retrieval in one release. ✓ Change one thing per rerun on the same cases, so you know what fixed it.
  • ✗ Treating every confident wrong answer as a hallucination. ✓ Check stop_reason, tool errors and cache hits; the model may have reasoned correctly from bad inputs.
  • ✗ Rolling back or blaming the new model on switch day. ✓ Replay frozen inputs on both models, apply the migration guide, fix instructions tuned for the old model, and decide on the numbers.
  • ✗ Repeating a buried rule in capitals. ✓ Resolve the conflict, make the examples show the rule, and enforce parsed formats with a schema.

Four reflexes, one method

A bigger modelsame inputs, same fault
A plea for accuracyadds no facts
Everything at onceno idea what worked
Roll back the modelthe cause stays
Reproduce, isolate, change one layerproven on the failing cases
A bigger model, a plea, a big-bang change and a rollback all leave the cause in place; finding the layer first and changing only it does not.

4.4.8 Put it together: diagnose three failures and prove each fix

You now have the whole method, the three families of failure, the lookalikes and the replay. The fastest way to make it stick is to cause each failure yourself, then fix it.

The rest of Domain 4 builds the machinery this lesson borrowed. Evaluation datasets (4.2) turn Mirela's 120 replayed questions into a standing regression set with deliberately chosen hard cases. A/B testing (4.3) takes a fix proven on a replay into live traffic. Token and cost optimisation (4.5) prices the max_tokens and effort choices, and monitoring (4.6) records the trace fields a replay needs on every request, before anyone complains.

Key takeaways

  • ✓ A bad answer is a symptom: reproduce it from the exact inputs, prompt version and model before changing anything.
  • ✓ Isolate the layer by swapping one at a time against the last known good version, and change one thing per rerun.
  • ✓ Prompt failures come from conflicting or buried instructions, missing context, open formats and misleading examples; fix the prompt, and enforce parsed formats with a schema.
  • ✓ For a hallucination, find where the fact went missing: not in the knowledge base, not retrieved, or in context but overridden; each needs a different fix.
  • ✓ Truncation, tool errors returned as empty results and stale caches look like model failures; stop_reason, tool results and cache hits expose them.
  • ✓ Prove model mismatch with a replay of frozen inputs on both models; fix settings and prompts tuned for the old model first, and move to a stronger tier only where the evidence shows a capability gap.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

30 CCAR-P questions on Domain 4, free

Every question in the bank is tagged to a domain, so you can drill 30 questions on Evaluation, Testing & Optimization alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 4 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources