Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 6 › Lesson 6.3

CCDV-F · Domain 6 · 11.0% of the exam · Lesson 6.3 · 21 min read

Output handling: structure it, check it, doubt it

How to get Claude's output in a fixed JSON shape, parse it defensively, check it against schema and business rules, and verify confident answers.

Written against skill 6.3 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

6.3.1 Why a receipt that parses can still be wrong

Every Friday, the staff of a small design agency photograph their receipts in an expense app, and the app turns each photo into an entry in the agency's accounting system. Sunniva, a back-end developer on the four-person team behind the app, wrote the feature that does the turning. Her code sends each photo to Claude with the instruction "Return the vendor, date, line items, subtotal, tax, total and VAT number as JSON". It parses the reply with json.loads and posts the result. In the demo it worked on every receipt she tried.

The first month with a real customer brought three failures. A long hotel bill came back as JSON that stopped halfway through its 70 line items, and the parser crashed the nightly job. A supermarket receipt parsed perfectly, but its line items added up to 47.30 while its total said 74.30, and the accounting system booked it anyway. The third was the quietest. Preparing a claim to recover VAT (value-added tax, the sales tax on receipts in many countries), Obinna, the agency's finance manager, found a taxi receipt with a VAT number the taxi firm had never printed. It was well formed, plausible and invented.

Only the first failure made any noise. The other two looked like success: valid JSON, every field present, nothing for a parser to catch. Claude's reply is text your code is about to act on, and nothing in how it looks tells you whether it is complete, consistent or true. Output handling is the work between the reply and anything that acts on it: the checks that decide whether a receipt reaches the ledger or goes to a person.

Three receipts, three failures

Hotel bill

JSON stops mid-liststop_reason: "max_tokens"
json.loads raisesthe nightly job halts

Supermarket receipt

Valid JSON, every field present
Items 47.30, total 74.30booked anyway

Taxi receipt

Valid JSON, every field present
A VAT number nobody printedwell formed and confident
Each receipt fails a different test (finished, consistent, true), and the only check Sunniva had, whether the text parsed, catches just the first.

6.3.2 Three ways to ask for a shape

The first question sounds too basic to matter: how do you get JSON your code can rely on at all? Sunniva asked for it in the prompt, and usually got it. But a prompt instruction is a request. Anthropic's documentation lists what can still come back: invalid JSON syntax, missing required fields and inconsistent data types. There are three ways to ask for structure, and they differ in what they guarantee. Two of them take a JSON schema, a standard JSON description of the fields, types and allowed values you expect.

Pattern How you ask What is guaranteed
Prompt only "Return JSON with these keys" in the prompt Nothing: prose, a missing field or a changed type can still arrive
Tool as a schema A tool such as record_receipt whose input_schema is the shape, with strict: true The tool's input matches the schema, but Claude still decides whether to call the tool
JSON outputs output_config.format with "type": "json_schema" and your schema The reply's text matches the schema, apart from refusals, max_tokens stops and enum capitalisation

Memorise the last column; the exceptions in the bottom row come back in the next section.

The middle row predates JSON outputs. You define a tool your code never runs; its input_schema is the shape you want, and strict: true holds the tool call's input to it. The catch is that Claude decides whether to call it. Older examples force the call with tool_choice, but on Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 a forced tool choice returns a 400 error. The documentation points instead to strict tools with the default auto setting, or to JSON outputs when the reply itself must have a fixed shape. A strict tool suits an action Claude may choose to take; when the whole reply IS the data, as with a receipt, JSON outputs are the direct route.

JSON outputs work by constrained decoding. Claude writes its reply one token at a time, a token being a short chunk of text, often part of a word. The API compiles your schema into a grammar that only lets Claude pick tokens that keep the reply valid. Think of a paper form with labelled boxes instead of a blank page: nobody can write an essay on it, but any box can still hold a wrong number. Below is Sunniva's schema as a model in Pydantic, the usual Python library for typed data; the SDK's transform_schema() turns it into a schema the API accepts. Look at currency, a closed set of values, and vat_number, which may be null and says when.

class LineItem(BaseModel):
    description: str
    amount: float

class Receipt(BaseModel):
    vendor: str
    date: date
    currency: Literal["EUR", "GBP", "USD"]                # a closed set, not free text
    line_items: list[LineItem]
    subtotal: float
    tax: float
    total: float
    vat_number: str | None = Field(description="As printed on the receipt; null if none is printed")

response = client.messages.create(
    model=MODEL, max_tokens=1024, messages=[receipt_message(photo)],   # photo plus instructions
    output_config={"format": {"type": "json_schema", "schema": transform_schema(Receipt)}},
)

Two limits of the schema language matter here. Every object needs additionalProperties: false, which the SDK adds for you. The API rejects numeric limits such as minimum with a 400 error, so transform_schema() moves them into the field's description, where they guide Claude but are enforced only when your code validates the reply. The SDK's client.messages.parse() does the transform, call and validation in one step; this lesson spells the steps out so you can see each check.

6.3.3 Defensive parsing: check the reply is finished before you read it

Sunniva switched to JSON outputs, and a week later another long hotel bill broke the job again. Here is the question that trips people up: if the API guarantees the schema, how can the JSON be broken? Because the guarantee has exceptions, and the documentation names them. Every response carries a stop_reason field that says why Claude stopped writing. When a reply hits your max_tokens cap on output tokens, it stops where it is, with stop_reason: "max_tokens", and the output may be incomplete and not match the schema. Her cap of 1,024 was plenty for a café receipt and far too little for 70 line items.

A refusal is the other exception. When Claude declines a request, the response is a normal HTTP 200 with stop_reason: "refusal", and its content list can be empty. Receipts rarely trigger one, but benign requests can trip the safety classifiers too. Code that reads response.content[0].text crashes on an empty list, and so does any code that assumes a text block exists. Streaming, where the reply arrives in pieces as it is written, adds a third hazard. The stop_reason arrives only in the message_delta event near the end, and a refusal can come mid-stream after partial output, which the documentation says to discard. A buffer that happens to parse as JSON halfway through is still not data.

Defensive parsing means your code assumes nothing until it has checked. Read stop_reason first: in a call like this one, with no tools or stop sequences, only end_turn means the reply finished. Find the text block by its type, not its position. Parse and validate the shape in one guarded step, and turn any failure into a held item with a reason, never a crash that halts the batch or a default value. In the code below, look at the stop_reason check, the search for the text block and the except branch; quarantine stands for your code that parks a receipt in a review queue.

def extract(receipt_id: str, response) -> Receipt | None:
    if response.stop_reason != "end_turn":            # max_tokens = cut off; refusal = no data
        quarantine(receipt_id, f"stop_reason={response.stop_reason}")
        return None
    text = next((b.text for b in response.content if b.type == "text"), None)
    if text is None:                                   # find the block by TYPE, never content[0]
        quarantine(receipt_id, "no text block")
        return None
    try:
        return Receipt.model_validate_json(text)       # PARSE and check the SHAPE in one step
    except ValidationError as err:                     # broken JSON, missing field, wrong type
        quarantine(receipt_id, str(err))               # a reason for the reviewer, never a default
        return None

It is tempting to rescue a truncated reply by closing its brackets. Resist it: the result is valid JSON with half the line items missing. For a max_tokens stop, the fix is more room: raise max_tokens or split the job, and run the receipt through the whole pipeline again. One more habit comes from the documentation. Structured outputs do not guarantee the capitalisation of enum values, so compare them case-insensitively; with Pydantic, normalise the case before the Literal check. Treat a value that still does not match as a failure, never as a cue for a default branch.

6.3.4 Validation: a valid shape is not valid data

The supermarket receipt would pass every check so far: complete JSON, every field present, and 74.30 is a perfectly good number. The schema can say the total is a number. It cannot say the total equals the subtotal plus the tax, because JSON Schema has no way to express rules across fields. Checking those is the job of response validation, and it has two layers.

The first layer is the schema check you already have: fields, types and allowed values. The second is business rules, the facts about your domain that no schema knows. For a receipt, the line items add up to the subtotal, and the subtotal plus the tax equals the total, each within a cent of rounding. The date is not in the future and falls inside the customer's expense-claim window. The code below adds those rules. Look at the one-cent tolerance and at the return value: an empty list is the only thing that lets a receipt through.

CENT = 0.01

def rule_problems(r: Receipt, today: date) -> list[str]:
    problems = []
    if abs(sum(item.amount for item in r.line_items) - r.subtotal) > CENT:
        problems.append("line items do not add up to the subtotal")
    if abs(r.subtotal + r.tax - r.total) > CENT:                   # tolerance absorbs rounding
        problems.append("subtotal plus tax does not equal the total")
    if r.date > today or (today - r.date).days > 90:               # the customer's claim window
        problems.append(f"date {r.date} is outside the claim window")
    return problems                                                # empty list = may be booked

What should happen when a rule fails? Not a quiet correction. Code that overwrites the total with the sum of the lines assumes the total was misread, and on the supermarket receipt it was a line item. Your code should fail closed: the receipt goes to review with its problems listed next to the photo, and a person settles it in seconds. Some teams add one bounded repair attempt first: send the problems back, ask Claude to read the receipt again, and run the new reply through every gate from the start.

The gates between Claude and the ledger

1. FINISHED?stop_reason is end_turn
2. SHAPEparse; fields, types, allowed values
3. RULEStotals add up, date in the window
4. VERIFYfacts that matter, against records
5. BOOKonly when every gate passed
A receipt reaches the accounting system only when every gate passes; a failure at any gate holds it for review with the reason attached.

6.3.5 Skepticism: confidence is not evidence

Now the taxi receipt, which no parser or arithmetic rule would ever catch. Why would Claude write a VAT number that is not there? A language model generates the most plausible continuation of what it has seen. When the output must contain a VAT number for a German taxi receipt, a well-formed German VAT number is a very plausible thing to write. Sunniva's first prompt asked for a VAT number and offered no way to say "none printed", so the model filled the box. An invented value arrives in the same format, with the same assured tone, as a value read off the paper.

Think of a fluent new hire who fills in every box on every form, because an empty box feels like failure. Their forms look perfect. Asking them to be more careful changes little; letting them leave a box empty, and checking the boxes that carry money against the records, changes a lot. Anthropic's guidance on reducing hallucinations (content that sounds right but is false or unsupported by the input) starts from the same place: explicitly give Claude permission to say it does not know. The nullable vat_number with its description does exactly that. The same guidance warns that such techniques reduce hallucinations without eliminating them, so critical information must still be validated.

So skepticism toward confident output is a design rule, not a mood: decide which facts carry real consequences, and check those against something other than the model. A confidence score you ask Claude for is generated the same way as the VAT number, and "are you sure?" gets an equally fluent answer. Running the same receipt several times and comparing is another technique the documentation suggests: disagreement can point to an invention, though agreement proves little. When the source is text rather than a photo, ask for the exact quote behind each claim and have your code confirm it appears in the source.

Match the check to the stakes. Obinna does not care whether a line reads "Espresso" or "Coffee", but a wrong VAT number costs the agency money and trouble with the tax office.

Field What rides on it How the pipeline checks it
Line item descriptions A label for the reader Accepted as read
Date The expense policy Business rule: not in the future, inside the claim window
Totals Money booked to the ledger Arithmetic in code, within a cent
VAT number The tax claim Null allowed; must match the vendor's record or an official registry lookup, otherwise a person checks it

Remember the principle rather than the rows: the higher the stakes, the more independent the check.

6.3.6 The exam traps

Every trap here trusts some property of the reply that does not prove it is safe to act on: that it was requested politely, that it parsed, that it sounded sure.

  • ✗ Asking for "JSON only" in the prompt and parsing whatever comes back. ✓ Send a schema with JSON outputs, or use a strict tool when the data is an action, so the shape is constrained rather than requested.
  • ✗ Treating a schema-valid reply as correct data. ✓ Check business rules in code. The schema guarantees the shape, with named exceptions, and never the arithmetic or the truth.
  • ✗ Parsing before reading stop_reason, or closing the brackets on a truncated reply. ✓ In an extraction call, only end_turn is a finished reply. Raise max_tokens or split the job; a patched array silently loses data.
  • ✗ Defaulting a missing or unexpected value so the pipeline keeps moving. ✓ Fail closed and hold the item with its reason. A default of 0, "approve" or "resolved" turns a detected error into an action.
  • ✗ Trusting a confidence score, an "are you sure?" follow-up or a bigger model to stop inventions. ✓ Allow "not present", and verify the facts that matter against a source you control.
  • ✗ Pulling fields out of a friendly explanation with a regular expression. ✓ Keep the machine-readable payload in its own schema. If people need an explanation too, give it a field of its own.

Four reasons to trust a reply, one that counts

"Return only JSON"a request, not a guarantee
It parsedcomplete and correct are other questions
confidence: 0.98more generated text
A bigger modelfewer mistakes, not none
Gates in your codefinished, shape, rules, verified, or held for review
A polite request, a clean parse, a confidence score and a bigger model all describe the reply; only checks in your own code decide whether it is safe to act on.

6.3.7 Put it together: build a receipt checker, then fool it

You now have every piece: an enforced schema, a parser that assumes nothing, business rules in code, verification for the facts that matter, and a review queue for the rest. The quickest way to make it stick is to build the pipeline on three receipts, then fool it on purpose and watch which gate catches what.

The same boundary appears across the rest of this guide. Debugging and error handling (4.1) covers what to do next with a held item: which failures deserve a retry, and how to tell a bug in your integration from a mistake in the output. AI application security (7.1) looks at the receipt itself as a threat, such as text in an uploaded photo that tries to give Claude instructions. Tool implementation (8.1) returns to strict schemas from the tool side, where a guaranteed shape still needs a sense check in your handler.

Key takeaways

  • ✓ Output handling is the set of checks between Claude's reply and anything that acts on it, because a reply that parses can still be incomplete, inconsistent or invented.
  • ✓ A prompt only requests a shape; JSON outputs (output_config.format) constrain the reply to your schema, and strict: true does the same for a tool's input.
  • ✓ The schema guarantee has named exceptions (refusals, max_tokens stops, enum capitalisation), so read stop_reason first and treat only end_turn as a finished extraction.
  • ✓ Parse defensively: find the text block by type, validate in one guarded step, and never patch truncated JSON or default a missing field.
  • ✓ A valid shape is not valid data: check business rules such as totals in your code, and hold anything that fails for review with its reason.
  • ✓ Confidence is not evidence: allow "not present", and verify high-stakes facts against your records or the source, never against the model's view of itself.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

18 CCDV-F questions on Domain 6, free

Every question in the bank is tagged to a domain, so you can drill 18 questions on Prompt and Context Engineering alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 6 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources