Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 4 › Lesson 4.1

CCDV-F · Domain 4 · 2.6% of the exam · Lesson 4.1 · 21 min read

Debugging and error handling: find where a failure starts

How to debug a Claude application: name the error type, pick the recovery that fits, read the trace, and tell an integration bug from a model mistake.

Written against skill 4.1 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

4.1.1 Why "the classifier is broken" is not a diagnosis

At 07:10 on a Tuesday, Noor opens her laptop to three messages. The overnight job that sorts customer feedback for an online homeware retailer logged thousands of errors. Hundreds of Claude's replies never reached the database because they would not parse. And the customer-care lead reports that the share of comments labelled praise jumped from 12% to 31% overnight, which nobody believes. Noor is on call, and the morning report goes out at nine.

The classifier is simple on paper. For each comment, a worker sends Claude the text with instructions and asks for a small JSON object: a category (delivery, product_quality, billing, website, praise or other), an urgent flag and a one-sentence reason. The worker parses the reply and stores the label. Yet each comment makes five hops on its way to the dashboard, and each hop can fail on its own.

One comment's path from source to dashboard

1. BUILDyour code: prompt, comment text, parameters
2. SENDthe HTTP call to the Messages API
3. GENERATEClaude writes the reply
4. PARSEyour code checks the JSON
5. STOREthe label the dashboard shows
A failure can start at any hop but only becomes visible at the last one, so the dashboard tells you that something broke, never where.

That is why the tempting first moves fail. A new prompt cannot fix a rate limit, a bigger model cannot fix a reply your own settings cut short, and a malformed request sent again only fails again. Any change made before you know where the failure started is a guess, and a guess made at 07:10 tends to add a second bug to the first.

The way out is to debug by origin. Hops 1, 2, 4 and 5 are the integration layer: your code that builds the request, the call to the API itself, and your code that handles the reply. Hop 3 is the model output: the text Claude wrote. To find the hop where a failure started, you read the trace, the record of what each hop did for one piece of work. Noor has four symptoms: bursts of 429 errors, bursts of 529 errors, replies that do not parse, and labels that parse but are wrong.

4.1.2 Name the failure before you touch it

Here is the question that decides everything that follows: what kind of failure is this? Failures come in two families. Loud failures raise an exception in your code: the API answered with an error status, or no answer arrived at all. Quiet failures arrive as HTTP 200, a normal success, and something is still wrong with the reply. For those, your clues are the reply itself and its stop_reason, the field in every successful response that says why Claude stopped writing: end_turn when it finished, max_tokens when it hit your cap.

For loud failures, the API has already done half the diagnosis. Every error response is JSON: an error object with a type and a message, and beside it a request_id. The same id comes back in the request-id header of every response, successful or not, and it is what Anthropic support asks for. Above all, the status code tells you who has to act. A 400, 401, 403, 404 or 413, a client error, says something on your side must change, so the same call sent again fails again; a 429, 500 or 529 says "not now", so waiting can work.

What you see What it means First move
400 invalid_request_error The format or content of your request is wrong Read the message; fix the request
401 authentication_error, 403 permission_error The key is bad, or lacks access to this resource Fix the key or its access; never retry in a loop
404 not_found_error, 413 request_too_large Wrong path or resource id; request over the size limit Fix the path or id; send less
429 rate_limit_error You went over a rate limit Wait as long as retry-after says, then slow down
500 api_error, 529 overloaded_error An internal error, or heavy load, on Anthropic's side Retry with backoff; keep the request_id
Timeout, connection error, 504 timeout_error No reply arrived, or none in time Retry a short call; stream a long one
200 with stop_reason max_tokens or refusal The reply was cut off at your cap, or declined Handle per stop reason; never parse a cut-off reply
200, but the reply does not parse Malformed output Validate, retry or repair, then find the cause
200, parses, and is wrong Wrong but well-formed output Isolate: the input you sent, or the model's judgment

Memorise the three-way split in the last column: fix, wait and retry, or check and validate. Recognise the individual codes well enough to place each one in it.

Quiet failures need more care, because nothing raises. The max_tokens cap you set is counted in tokens (the word pieces a model reads and writes), and a reply that hits it can end mid-JSON while the API reports success, so only your code can notice. A malformed reply fails to parse or misses a field. Worst of all is output that parses, passes every check and is wrong; only a comparison with reality catches it.

Agents add two more kinds, around tools. A tool error means a tool your code runs has failed, such as an order lookup that timed out. An invalid tool call means Claude asked for a tool with a missing or wrong parameter. The first starts in your integration, the second in the model's output. Sorted this way, Noor's night looks less chaotic: the 429s and 529s are loud and possibly transient, while the broken replies and wrong labels are quiet and need a closer look.

4.1.3 Match the recovery to the error type

Once the error type is known, the recovery almost picks itself, because there are only four to choose from.

Four recoveries, each for its own error types

RETRY

429, 500, 529and timeouts on short calls
Backoff, jitter, a budgethonour retry-after

FIX THE REQUEST

400, 401, 403, 404, 413and max_tokens cut-offs
Change code or confignever resend it unchanged

VALIDATE AND REPAIR

Malformed outputa 200 that will not parse
One retry, or flag itnever store half-valid data

FALL BACK

Retries ran outor the path is down
Queue, default, persondegrade, never drop
Retrying only helps when time is the cure; fixing, validating and falling back cover everything else.

Retry is for errors where time is the cure, and the skill is in how you wait. Picture a busy phone line. Redial instantly fifty times and you keep the line busy yourself; have everyone redial on the hour and it jams again on the hour. So you wait longer after each failure (exponential backoff: 1, 2, 4, 8 seconds, up to a ceiling). You add a small random amount to every wait (jitter), so that many workers do not come back at the same instant. When a 429 carries a retry-after header, you wait that many seconds, because earlier retries will fail. And you stop after a fixed number of attempts, because retrying without end is an outage that costs money.

Fix the request covers everything that fails the same way every time. A 400 names the faulty part of the request in its message. A reply cut off at max_tokens belongs here too, because the same cap cuts the same reply again. Validate and repair is for malformed output: check every reply against the shape you expect, retry once or mark the item for review, and never pass half-valid data on. Fall back is what happens when nothing else helps in time: put the item back in a queue, return a safe default such as "unclassified", or hand it to a person. The job degrades instead of crashing or silently losing data.

The official SDKs already retry connection errors, 429s and 5xx errors twice by default, with exponential backoff that honours retry-after. When you need your own policy, such as a longer budget with a fallback, write a wrapper like this one and turn the SDK's retries off. Retry layers multiply: three SDK attempts inside five of yours make fifteen calls for one comment. Look at the if that refuses to retry client errors, the line that prefers the server's retry-after, and the random extra that adds jitter.

import random, time
import anthropic

client = anthropic.Anthropic(max_retries=0)    # this wrapper owns the retries

def create_with_recovery(request: dict, attempts: int = 5):
    for attempt in range(attempts):
        try:
            return client.messages.create(**request)
        except anthropic.APIStatusError as err:
            if err.status_code != 429 and err.status_code < 500:
                raise                          # 400, 401, 403, 404, 413: FIX the request
            hint = err.response.headers.get("retry-after")   # sent with a normal 429
        except anthropic.APIConnectionError:   # network drop or client-side timeout
            hint = None
        delay = float(hint) if hint else min(60, 2 ** attempt)   # else 1, 2, 4, 8 ... s
        time.sleep(delay + random.uniform(0, delay / 2))         # JITTER spreads the workers
    return None                                # budget spent: FALL BACK, e.g. requeue it

Two exceptions keep this honest. Not every 429 is temporary. When an organisation reaches its usage tier's monthly spend cap, requests return a 429 rate_limit_error with no retry-after header and an error_code of enforced_spend_limit_reached in the error details. Every retry then fails until access resumes. And a retry is safe only when repeating the call does no harm. A classification is fine, but a step that creates something, such as a refund, first needs an idempotency key: a unique id per operation that lets the receiving service ignore a repeat.

Noor's 429s are partly her own team's doing. Callum's Monday release raised the worker count from 8 to 64 to clear a backlog, and the old retry code resent every failure at once. A sharp jump in usage can trip Anthropic's acceleration limits, and a per-minute limit may be enforced over shorter intervals, so bursts fail even when the minute's total looks fine. Her fix is backoff with jitter, honouring retry-after, and a gradual ramp-up. The 529s come from heavy traffic across the whole API, so backoff alone is the answer, with a requeue for any comment still failing after five attempts.

4.1.4 Read the trace, not the dashboard

The dashboard tells Noor what went wrong in aggregate; the traces tell her why, one piece of work at a time. Think of a parcel's tracking history. When a parcel goes missing, nobody searches the whole country; they find the last scan that was fine and the first one that was not.

A trace can only answer questions you recorded answers to. For each model call, log a trace id that ties it to the comment, plus the request_id, the prompt version, the model and key parameters such as max_tokens. Add each attempt and its status, the stop_reason, token usage, the raw reply text and the parse result. The Python and TypeScript SDKs expose the request id as _request_id on each response object. The Claude Agent SDK can also export traces (in beta) in OpenTelemetry, the open tracing standard, with a span, one step's record, for each model request and tool call.

Here is one of Noor's failing traces. Read three fields together: stop_reason, the output token count, and the raw reply that ends mid-sentence. They show a reply cut off at the 40-token cap, not a garbled one.

{
  "trace_id": "fb-0929-018812",
  "prompt_version": "classify-v7",
  "model": "claude-haiku-4-5",
  "max_tokens": 40,
  "input_chars": 312,
  "attempts": [
    {"status": 529, "request_id": "req_011CT...", "waited_s": 1.4},
    {"status": 200, "request_id": "req_011CT..."}
  ],
  "stop_reason": "max_tokens",
  "usage": {"input_tokens": 431, "output_tokens": 40},
  "raw_reply": "{\"category\": \"delivery\", \"urgent\": true, \"reason\": \"Sofa arrived two weeks late with a torn cover and a missing leg, and the customer says",
  "parse": "JSONDecodeError: Unterminated string",
  "stored_label": null
}

One trace shows one failure. A failure mode is a pattern: many traces failing the same way for the same reason. To find them, group the failing traces by signature (status code, stop_reason, parse error, disputed label). Then ask two questions of each group. What do all its traces share that the successful ones lack? And when did the first one appear, next to which change?

Signature in the traces What the failing traces share Origin
429 with retry-after, resent at once Began at 02:00, when 64 workers started together Integration: concurrency and retry policy
529 across all workers 02:40 to 03:10, with nothing changed on the team's side Integration: the API itself, under heavy load (transient)
200, stop_reason max_tokens, parse error Always exactly 40 output tokens; began with prompt v7 Integration: max_tokens too low for the new reason field
Parses, wrong label, input_chars 500 Every comment longer than 500 characters Integration: a new helper trims comments
Parses, wrong label, full input Sarcastic comments ("Great, a third late delivery") Model output

Grouping turned thousands of failures into five failure modes. Three of them trace back to Monday's release: the worker count, prompt v7's new reason field with the old max_tokens, and a helper that trims comments to save tokens. The sarcasm row is older. Those comments had been mislabelled for weeks, and the spike made someone finally look.

4.1.5 Integration layer or model output?

Two rows in Noor's table look identical from the dashboard: a label that parses and is wrong. One comes from her team's code, the other from Claude. Telling them apart is the heart of this skill: check three points along the path, in order, and stop at the first one that fails.

Three checkpoints, in order

1. THE REQUESTdid it carry what you intended?
2. THE REPLYgiven that exact request, is the raw text right?
3. THE HANDLINGdid your code turn a right reply into the right result?
The first checkpoint that fails is the origin. A wrong request or wrong handling is an integration bug; a right request with a wrong reply is a model-output problem.

At the first checkpoint, compare the request in the trace with the source data: the right text, the whole text, the right prompt version, sensible parameters. For the 500-character group, Noor opens comment 18,390. The customer spends three sentences praising years of good service before the complaint begins, and the request stops after the praise. Claude labelled exactly what it was given. The fault is in the integration layer, and no prompt change could ever reach it.

At the second checkpoint the request is right, so you test the reply. Replay the request exactly as the trace recorded it and read the raw text before any parsing. Replay more than once, because a model can answer differently from run to run. Then vary one thing at a time. When Noor replays the trimmed comment with its full text, the label comes back delivery, which confirms the origin. When she replays a sarcastic comment five times with its full text, it comes back praise every time. That is a model-output problem, and the fix lives in the prompt: a note on sarcasm and a few labelled examples, checked against known comments before it ships.

The third checkpoint catches bugs that happen after a correct reply, such as mapping labels to the wrong column or storing a result under the wrong comment. There the raw reply in the trace is right and the stored result is not, so the fault is in your code again.

The same order works for each tool call in an agent. If the trace shows the right tool called with sensible inputs and the service behind it failed, the model did its part. Recover inside the tool code with the same bounded backoff; if that fails too, return a tool_result with is_error: true and a message saying what went wrong, so Claude can adapt. Never rerun the whole conversation to recover one call. If the inputs themselves are wrong or missing, the origin is the model's output, and a clearer tool description is the usual fix.

4.1.6 The exam traps

Every trap here is a fix applied before the failure was classified or located, so it feels like progress and skips the diagnosis.

  • ✗ Retrying every error the same way. ✓ Retry only transient errors (429, 500, 529, timeouts), with a fixed budget. A 400 or a 401 fails the same way every time, so fix the request or the key.
  • ✗ Retrying at once, with every worker together. ✓ Back off exponentially, add jitter and honour retry-after. Instant retries turn a brief limit into a long outage of your own making.
  • ✗ Treating HTTP 200 as success. ✓ Check stop_reason and validate the content. A cut-off, a malformed reply and a wrong label all arrive with status 200.
  • ✗ Rewording the prompt or switching models before reading the trace. ✓ Isolate the origin first. A prompt change cannot fix a trimmed input, and a bigger model cannot lift a max_tokens cap.
  • ✗ Rerunning the whole job or conversation to recover one step. ✓ Retry or replay only the step that failed, with the inputs its trace recorded. Full reruns cost time and money and can repeat side effects.
  • ✗ Logging only errors and final results. ✓ Log every call, with its request, parameters, stop_reason and raw reply. Without them, quiet failures leave no trail.

4.1.7 Put it together: debug a classifier you broke on purpose

You now have every piece: the error types, the recovery that fits each, traces, and three checkpoints that separate integration bugs from model output. The fastest way to make them stick is to build a small classifier, break it the ways Noor's broke, and read your own traces.

Three later skills build on this one. Technical fundamentals (5.2) covers the SDKs' retry and timeout settings. Output handling (6.3) turns "validate and repair" into concrete patterns, and tool implementation (8.1) shapes what a failing tool returns so Claude can recover inside the loop.

Key takeaways

  • ✓ A failure starts either in the integration layer or in the model's output; find which one before you change anything.
  • ✓ Identify the error type: a 400, 401, 403, 404 or 413 means fix the request; 429, 500, 529 and timeouts are transient; a 200 still needs its stop_reason checked and its content validated.
  • ✓ Retry only transient errors, with exponential backoff, jitter, retry-after honoured and a fixed budget; a 429 from the monthly spend cap is not transient.
  • ✓ When a retry cannot help, fix the request, validate and repair the output, or fall back to a queue, a default or a person.
  • ✓ A trace records every step of one piece of work; group failing traces by signature and compare them with successful ones and with recent changes.
  • ✓ Isolate with three checkpoints (the request, the raw reply, your handling), and replay recorded requests to tell an integration bug from a model-output problem.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

3 CCDV-F questions on Domain 4, free

Every question in the bank is tagged to a domain, so you can drill 3 questions on Eval, Testing, and Debugging alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 4 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources