Home › Study guides › CCDV-F › Domain 4 › Lesson 4.1
CCDV-F · Domain 4 · 2.6% of the exam · Lesson 4.1 · 21 min read
Debugging and error handling: find where a failure starts
How to debug a Claude application: name the error type, pick the recovery that fits, read the trace, and tell an integration bug from a model mistake.
Written against skill 4.1 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.1.1 Why "the classifier is broken" is not a diagnosis
At 07:10 on a Tuesday, Noor opens her laptop to three messages. The overnight job that sorts customer feedback for an online homeware retailer logged thousands of errors. Hundreds of Claude's replies never reached the database because they would not parse. And the customer-care lead reports that the share of comments labelled praise jumped from 12% to 31% overnight, which nobody believes. Noor is on call, and the morning report goes out at nine.
The classifier is simple on paper. For each comment, a worker sends Claude the text with instructions and asks for a small JSON object: a category (delivery, product_quality, billing, website, praise or other), an urgent flag and a one-sentence reason. The worker parses the reply and stores the label. Yet each comment makes five hops on its way to the dashboard, and each hop can fail on its own.
One comment's path from source to dashboard
That is why the tempting first moves fail. A new prompt cannot fix a rate limit, a bigger model cannot fix a reply your own settings cut short, and a malformed request sent again only fails again. Any change made before you know where the failure started is a guess, and a guess made at 07:10 tends to add a second bug to the first.
The way out is to debug by origin. Hops 1, 2, 4 and 5 are the integration layer: your code that builds the request, the call to the API itself, and your code that handles the reply. Hop 3 is the model output: the text Claude wrote. To find the hop where a failure started, you read the trace, the record of what each hop did for one piece of work. Noor has four symptoms: bursts of 429 errors, bursts of 529 errors, replies that do not parse, and labels that parse but are wrong.
4.1.2 Name the failure before you touch it
Here is the question that decides everything that follows: what kind of failure is this? Failures come in two families. Loud failures raise an exception in your code: the API answered with an error status, or no answer arrived at all. Quiet failures arrive as HTTP 200, a normal success, and something is still wrong with the reply. For those, your clues are the reply itself and its stop_reason, the field in every successful response that says why Claude stopped writing: end_turn when it finished, max_tokens when it hit your cap.
For loud failures, the API has already done half the diagnosis. Every error response is JSON: an error object with a type and a message, and beside it a request_id. The same id comes back in the request-id header of every response, successful or not, and it is what Anthropic support asks for. Above all, the status code tells you who has to act. A 400, 401, 403, 404 or 413, a client error, says something on your side must change, so the same call sent again fails again; a 429, 500 or 529 says "not now", so waiting can work.
| What you see | What it means | First move |
|---|---|---|
400 invalid_request_error |
The format or content of your request is wrong | Read the message; fix the request |
401 authentication_error, 403 permission_error |
The key is bad, or lacks access to this resource | Fix the key or its access; never retry in a loop |
404 not_found_error, 413 request_too_large |
Wrong path or resource id; request over the size limit | Fix the path or id; send less |
429 rate_limit_error |
You went over a rate limit | Wait as long as retry-after says, then slow down |
500 api_error, 529 overloaded_error |
An internal error, or heavy load, on Anthropic's side | Retry with backoff; keep the request_id |
Timeout, connection error, 504 timeout_error |
No reply arrived, or none in time | Retry a short call; stream a long one |
200 with stop_reason max_tokens or refusal |
The reply was cut off at your cap, or declined | Handle per stop reason; never parse a cut-off reply |
| 200, but the reply does not parse | Malformed output | Validate, retry or repair, then find the cause |
| 200, parses, and is wrong | Wrong but well-formed output | Isolate: the input you sent, or the model's judgment |
Memorise the three-way split in the last column: fix, wait and retry, or check and validate. Recognise the individual codes well enough to place each one in it.
Quiet failures need more care, because nothing raises. The max_tokens cap you set is counted in tokens (the word pieces a model reads and writes), and a reply that hits it can end mid-JSON while the API reports success, so only your code can notice. A malformed reply fails to parse or misses a field. Worst of all is output that parses, passes every check and is wrong; only a comparison with reality catches it.
Agents add two more kinds, around tools. A tool error means a tool your code runs has failed, such as an order lookup that timed out. An invalid tool call means Claude asked for a tool with a missing or wrong parameter. The first starts in your integration, the second in the model's output. Sorted this way, Noor's night looks less chaotic: the 429s and 529s are loud and possibly transient, while the broken replies and wrong labels are quiet and need a closer look.
4.1.3 Match the recovery to the error type
Once the error type is known, the recovery almost picks itself, because there are only four to choose from.
Four recoveries, each for its own error types
RETRY
retry-afterFIX THE REQUEST
max_tokens cut-offsVALIDATE AND REPAIR
FALL BACK
Retry is for errors where time is the cure, and the skill is in how you wait. Picture a busy phone line. Redial instantly fifty times and you keep the line busy yourself; have everyone redial on the hour and it jams again on the hour. So you wait longer after each failure (exponential backoff: 1, 2, 4, 8 seconds, up to a ceiling). You add a small random amount to every wait (jitter), so that many workers do not come back at the same instant. When a 429 carries a retry-after header, you wait that many seconds, because earlier retries will fail. And you stop after a fixed number of attempts, because retrying without end is an outage that costs money.
Fix the request covers everything that fails the same way every time. A 400 names the faulty part of the request in its message. A reply cut off at max_tokens belongs here too, because the same cap cuts the same reply again. Validate and repair is for malformed output: check every reply against the shape you expect, retry once or mark the item for review, and never pass half-valid data on. Fall back is what happens when nothing else helps in time: put the item back in a queue, return a safe default such as "unclassified", or hand it to a person. The job degrades instead of crashing or silently losing data.
The official SDKs already retry connection errors, 429s and 5xx errors twice by default, with exponential backoff that honours retry-after. When you need your own policy, such as a longer budget with a fallback, write a wrapper like this one and turn the SDK's retries off. Retry layers multiply: three SDK attempts inside five of yours make fifteen calls for one comment. Look at the if that refuses to retry client errors, the line that prefers the server's retry-after, and the random extra that adds jitter.
import random, time
import anthropic
client = anthropic.Anthropic(max_retries=0) # this wrapper owns the retries
def create_with_recovery(request: dict, attempts: int = 5):
for attempt in range(attempts):
try:
return client.messages.create(**request)
except anthropic.APIStatusError as err:
if err.status_code != 429 and err.status_code < 500:
raise # 400, 401, 403, 404, 413: FIX the request
hint = err.response.headers.get("retry-after") # sent with a normal 429
except anthropic.APIConnectionError: # network drop or client-side timeout
hint = None
delay = float(hint) if hint else min(60, 2 ** attempt) # else 1, 2, 4, 8 ... s
time.sleep(delay + random.uniform(0, delay / 2)) # JITTER spreads the workers
return None # budget spent: FALL BACK, e.g. requeue it
Two exceptions keep this honest. Not every 429 is temporary. When an organisation reaches its usage tier's monthly spend cap, requests return a 429 rate_limit_error with no retry-after header and an error_code of enforced_spend_limit_reached in the error details. Every retry then fails until access resumes. And a retry is safe only when repeating the call does no harm. A classification is fine, but a step that creates something, such as a refund, first needs an idempotency key: a unique id per operation that lets the receiving service ignore a repeat.
Noor's 429s are partly her own team's doing. Callum's Monday release raised the worker count from 8 to 64 to clear a backlog, and the old retry code resent every failure at once. A sharp jump in usage can trip Anthropic's acceleration limits, and a per-minute limit may be enforced over shorter intervals, so bursts fail even when the minute's total looks fine. Her fix is backoff with jitter, honouring retry-after, and a gradual ramp-up. The 529s come from heavy traffic across the whole API, so backoff alone is the answer, with a requeue for any comment still failing after five attempts.
4.1.4 Read the trace, not the dashboard
The dashboard tells Noor what went wrong in aggregate; the traces tell her why, one piece of work at a time. Think of a parcel's tracking history. When a parcel goes missing, nobody searches the whole country; they find the last scan that was fine and the first one that was not.
A trace can only answer questions you recorded answers to. For each model call, log a trace id that ties it to the comment, plus the request_id, the prompt version, the model and key parameters such as max_tokens. Add each attempt and its status, the stop_reason, token usage, the raw reply text and the parse result. The Python and TypeScript SDKs expose the request id as _request_id on each response object. The Claude Agent SDK can also export traces (in beta) in OpenTelemetry, the open tracing standard, with a span, one step's record, for each model request and tool call.
Here is one of Noor's failing traces. Read three fields together: stop_reason, the output token count, and the raw reply that ends mid-sentence. They show a reply cut off at the 40-token cap, not a garbled one.
{
"trace_id": "fb-0929-018812",
"prompt_version": "classify-v7",
"model": "claude-haiku-4-5",
"max_tokens": 40,
"input_chars": 312,
"attempts": [
{"status": 529, "request_id": "req_011CT...", "waited_s": 1.4},
{"status": 200, "request_id": "req_011CT..."}
],
"stop_reason": "max_tokens",
"usage": {"input_tokens": 431, "output_tokens": 40},
"raw_reply": "{\"category\": \"delivery\", \"urgent\": true, \"reason\": \"Sofa arrived two weeks late with a torn cover and a missing leg, and the customer says",
"parse": "JSONDecodeError: Unterminated string",
"stored_label": null
}
One trace shows one failure. A failure mode is a pattern: many traces failing the same way for the same reason. To find them, group the failing traces by signature (status code, stop_reason, parse error, disputed label). Then ask two questions of each group. What do all its traces share that the successful ones lack? And when did the first one appear, next to which change?
| Signature in the traces | What the failing traces share | Origin |
|---|---|---|
429 with retry-after, resent at once |
Began at 02:00, when 64 workers started together | Integration: concurrency and retry policy |
| 529 across all workers | 02:40 to 03:10, with nothing changed on the team's side | Integration: the API itself, under heavy load (transient) |
200, stop_reason max_tokens, parse error |
Always exactly 40 output tokens; began with prompt v7 | Integration: max_tokens too low for the new reason field |
Parses, wrong label, input_chars 500 |
Every comment longer than 500 characters | Integration: a new helper trims comments |
| Parses, wrong label, full input | Sarcastic comments ("Great, a third late delivery") | Model output |
Grouping turned thousands of failures into five failure modes. Three of them trace back to Monday's release: the worker count, prompt v7's new reason field with the old max_tokens, and a helper that trims comments to save tokens. The sarcasm row is older. Those comments had been mislabelled for weeks, and the spike made someone finally look.
4.1.5 Integration layer or model output?
Two rows in Noor's table look identical from the dashboard: a label that parses and is wrong. One comes from her team's code, the other from Claude. Telling them apart is the heart of this skill: check three points along the path, in order, and stop at the first one that fails.
Three checkpoints, in order
At the first checkpoint, compare the request in the trace with the source data: the right text, the whole text, the right prompt version, sensible parameters. For the 500-character group, Noor opens comment 18,390. The customer spends three sentences praising years of good service before the complaint begins, and the request stops after the praise. Claude labelled exactly what it was given. The fault is in the integration layer, and no prompt change could ever reach it.
At the second checkpoint the request is right, so you test the reply. Replay the request exactly as the trace recorded it and read the raw text before any parsing. Replay more than once, because a model can answer differently from run to run. Then vary one thing at a time. When Noor replays the trimmed comment with its full text, the label comes back delivery, which confirms the origin. When she replays a sarcastic comment five times with its full text, it comes back praise every time. That is a model-output problem, and the fix lives in the prompt: a note on sarcasm and a few labelled examples, checked against known comments before it ships.
The third checkpoint catches bugs that happen after a correct reply, such as mapping labels to the wrong column or storing a result under the wrong comment. There the raw reply in the trace is right and the stored result is not, so the fault is in your code again.
The same order works for each tool call in an agent. If the trace shows the right tool called with sensible inputs and the service behind it failed, the model did its part. Recover inside the tool code with the same bounded backoff; if that fails too, return a tool_result with is_error: true and a message saying what went wrong, so Claude can adapt. Never rerun the whole conversation to recover one call. If the inputs themselves are wrong or missing, the origin is the model's output, and a clearer tool description is the usual fix.
4.1.6 The exam traps
Every trap here is a fix applied before the failure was classified or located, so it feels like progress and skips the diagnosis.
- ✗ Retrying every error the same way. ✓ Retry only transient errors (429, 500, 529, timeouts), with a fixed budget. A 400 or a 401 fails the same way every time, so fix the request or the key.
- ✗ Retrying at once, with every worker together. ✓ Back off exponentially, add jitter and honour
retry-after. Instant retries turn a brief limit into a long outage of your own making. - ✗ Treating HTTP 200 as success. ✓ Check
stop_reasonand validate the content. A cut-off, a malformed reply and a wrong label all arrive with status 200. - ✗ Rewording the prompt or switching models before reading the trace. ✓ Isolate the origin first. A prompt change cannot fix a trimmed input, and a bigger model cannot lift a
max_tokenscap. - ✗ Rerunning the whole job or conversation to recover one step. ✓ Retry or replay only the step that failed, with the inputs its trace recorded. Full reruns cost time and money and can repeat side effects.
- ✗ Logging only errors and final results. ✓ Log every call, with its request, parameters,
stop_reasonand raw reply. Without them, quiet failures leave no trail.
4.1.7 Put it together: debug a classifier you broke on purpose
You now have every piece: the error types, the recovery that fits each, traces, and three checkpoints that separate integration bugs from model output. The fastest way to make them stick is to build a small classifier, break it the ways Noor's broke, and read your own traces.
Three later skills build on this one. Technical fundamentals (5.2) covers the SDKs' retry and timeout settings. Output handling (6.3) turns "validate and repair" into concrete patterns, and tool implementation (8.1) shapes what a failing tool returns so Claude can recover inside the loop.
Key takeaways
- ✓ A failure starts either in the integration layer or in the model's output; find which one before you change anything.
- ✓ Identify the error type: a 400, 401, 403, 404 or 413 means fix the request; 429, 500, 529 and timeouts are transient; a 200 still needs its
stop_reasonchecked and its content validated. - ✓ Retry only transient errors, with exponential backoff, jitter,
retry-afterhonoured and a fixed budget; a 429 from the monthly spend cap is not transient. - ✓ When a retry cannot help, fix the request, validate and repair the output, or fall back to a queue, a default or a person.
- ✓ A trace records every step of one piece of work; group failing traces by signature and compare them with successful ones and with recent changes.
- ✓ Isolate with three checkpoints (the request, the raw reply, your handling), and replay recorded requests to tell an integration bug from a model-output problem.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
3 CCDV-F questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 3 questions on Eval, Testing, and Debugging alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.