Home › Study guides › CCDV-F › Domain 6 › Lesson 6.3
CCDV-F · Domain 6 · 11.0% of the exam · Lesson 6.3 · 21 min read
Output handling: structure it, check it, doubt it
How to get Claude's output in a fixed JSON shape, parse it defensively, check it against schema and business rules, and verify confident answers.
Written against skill 6.3 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
6.3.1 Why a receipt that parses can still be wrong
Every Friday, the staff of a small design agency photograph their receipts in an expense app, and the app turns each photo into an entry in the agency's accounting system. Sunniva, a back-end developer on the four-person team behind the app, wrote the feature that does the turning. Her code sends each photo to Claude with the instruction "Return the vendor, date, line items, subtotal, tax, total and VAT number as JSON". It parses the reply with json.loads and posts the result. In the demo it worked on every receipt she tried.
The first month with a real customer brought three failures. A long hotel bill came back as JSON that stopped halfway through its 70 line items, and the parser crashed the nightly job. A supermarket receipt parsed perfectly, but its line items added up to 47.30 while its total said 74.30, and the accounting system booked it anyway. The third was the quietest. Preparing a claim to recover VAT (value-added tax, the sales tax on receipts in many countries), Obinna, the agency's finance manager, found a taxi receipt with a VAT number the taxi firm had never printed. It was well formed, plausible and invented.
Only the first failure made any noise. The other two looked like success: valid JSON, every field present, nothing for a parser to catch. Claude's reply is text your code is about to act on, and nothing in how it looks tells you whether it is complete, consistent or true. Output handling is the work between the reply and anything that acts on it: the checks that decide whether a receipt reaches the ledger or goes to a person.
Three receipts, three failures
Hotel bill
stop_reason: "max_tokens"json.loads raisesthe nightly job haltsSupermarket receipt
Taxi receipt
6.3.2 Three ways to ask for a shape
The first question sounds too basic to matter: how do you get JSON your code can rely on at all? Sunniva asked for it in the prompt, and usually got it. But a prompt instruction is a request. Anthropic's documentation lists what can still come back: invalid JSON syntax, missing required fields and inconsistent data types. There are three ways to ask for structure, and they differ in what they guarantee. Two of them take a JSON schema, a standard JSON description of the fields, types and allowed values you expect.
| Pattern | How you ask | What is guaranteed |
|---|---|---|
| Prompt only | "Return JSON with these keys" in the prompt | Nothing: prose, a missing field or a changed type can still arrive |
| Tool as a schema | A tool such as record_receipt whose input_schema is the shape, with strict: true |
The tool's input matches the schema, but Claude still decides whether to call the tool |
| JSON outputs | output_config.format with "type": "json_schema" and your schema |
The reply's text matches the schema, apart from refusals, max_tokens stops and enum capitalisation |
Memorise the last column; the exceptions in the bottom row come back in the next section.
The middle row predates JSON outputs. You define a tool your code never runs; its input_schema is the shape you want, and strict: true holds the tool call's input to it. The catch is that Claude decides whether to call it. Older examples force the call with tool_choice, but on Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 a forced tool choice returns a 400 error. The documentation points instead to strict tools with the default auto setting, or to JSON outputs when the reply itself must have a fixed shape. A strict tool suits an action Claude may choose to take; when the whole reply IS the data, as with a receipt, JSON outputs are the direct route.
JSON outputs work by constrained decoding. Claude writes its reply one token at a time, a token being a short chunk of text, often part of a word. The API compiles your schema into a grammar that only lets Claude pick tokens that keep the reply valid. Think of a paper form with labelled boxes instead of a blank page: nobody can write an essay on it, but any box can still hold a wrong number. Below is Sunniva's schema as a model in Pydantic, the usual Python library for typed data; the SDK's transform_schema() turns it into a schema the API accepts. Look at currency, a closed set of values, and vat_number, which may be null and says when.
class LineItem(BaseModel):
description: str
amount: float
class Receipt(BaseModel):
vendor: str
date: date
currency: Literal["EUR", "GBP", "USD"] # a closed set, not free text
line_items: list[LineItem]
subtotal: float
tax: float
total: float
vat_number: str | None = Field(description="As printed on the receipt; null if none is printed")
response = client.messages.create(
model=MODEL, max_tokens=1024, messages=[receipt_message(photo)], # photo plus instructions
output_config={"format": {"type": "json_schema", "schema": transform_schema(Receipt)}},
)
Two limits of the schema language matter here. Every object needs additionalProperties: false, which the SDK adds for you. The API rejects numeric limits such as minimum with a 400 error, so transform_schema() moves them into the field's description, where they guide Claude but are enforced only when your code validates the reply. The SDK's client.messages.parse() does the transform, call and validation in one step; this lesson spells the steps out so you can see each check.
6.3.3 Defensive parsing: check the reply is finished before you read it
Sunniva switched to JSON outputs, and a week later another long hotel bill broke the job again. Here is the question that trips people up: if the API guarantees the schema, how can the JSON be broken? Because the guarantee has exceptions, and the documentation names them. Every response carries a stop_reason field that says why Claude stopped writing. When a reply hits your max_tokens cap on output tokens, it stops where it is, with stop_reason: "max_tokens", and the output may be incomplete and not match the schema. Her cap of 1,024 was plenty for a café receipt and far too little for 70 line items.
A refusal is the other exception. When Claude declines a request, the response is a normal HTTP 200 with stop_reason: "refusal", and its content list can be empty. Receipts rarely trigger one, but benign requests can trip the safety classifiers too. Code that reads response.content[0].text crashes on an empty list, and so does any code that assumes a text block exists. Streaming, where the reply arrives in pieces as it is written, adds a third hazard. The stop_reason arrives only in the message_delta event near the end, and a refusal can come mid-stream after partial output, which the documentation says to discard. A buffer that happens to parse as JSON halfway through is still not data.
Defensive parsing means your code assumes nothing until it has checked. Read stop_reason first: in a call like this one, with no tools or stop sequences, only end_turn means the reply finished. Find the text block by its type, not its position. Parse and validate the shape in one guarded step, and turn any failure into a held item with a reason, never a crash that halts the batch or a default value. In the code below, look at the stop_reason check, the search for the text block and the except branch; quarantine stands for your code that parks a receipt in a review queue.
def extract(receipt_id: str, response) -> Receipt | None:
if response.stop_reason != "end_turn": # max_tokens = cut off; refusal = no data
quarantine(receipt_id, f"stop_reason={response.stop_reason}")
return None
text = next((b.text for b in response.content if b.type == "text"), None)
if text is None: # find the block by TYPE, never content[0]
quarantine(receipt_id, "no text block")
return None
try:
return Receipt.model_validate_json(text) # PARSE and check the SHAPE in one step
except ValidationError as err: # broken JSON, missing field, wrong type
quarantine(receipt_id, str(err)) # a reason for the reviewer, never a default
return None
It is tempting to rescue a truncated reply by closing its brackets. Resist it: the result is valid JSON with half the line items missing. For a max_tokens stop, the fix is more room: raise max_tokens or split the job, and run the receipt through the whole pipeline again. One more habit comes from the documentation. Structured outputs do not guarantee the capitalisation of enum values, so compare them case-insensitively; with Pydantic, normalise the case before the Literal check. Treat a value that still does not match as a failure, never as a cue for a default branch.
6.3.4 Validation: a valid shape is not valid data
The supermarket receipt would pass every check so far: complete JSON, every field present, and 74.30 is a perfectly good number. The schema can say the total is a number. It cannot say the total equals the subtotal plus the tax, because JSON Schema has no way to express rules across fields. Checking those is the job of response validation, and it has two layers.
The first layer is the schema check you already have: fields, types and allowed values. The second is business rules, the facts about your domain that no schema knows. For a receipt, the line items add up to the subtotal, and the subtotal plus the tax equals the total, each within a cent of rounding. The date is not in the future and falls inside the customer's expense-claim window. The code below adds those rules. Look at the one-cent tolerance and at the return value: an empty list is the only thing that lets a receipt through.
CENT = 0.01
def rule_problems(r: Receipt, today: date) -> list[str]:
problems = []
if abs(sum(item.amount for item in r.line_items) - r.subtotal) > CENT:
problems.append("line items do not add up to the subtotal")
if abs(r.subtotal + r.tax - r.total) > CENT: # tolerance absorbs rounding
problems.append("subtotal plus tax does not equal the total")
if r.date > today or (today - r.date).days > 90: # the customer's claim window
problems.append(f"date {r.date} is outside the claim window")
return problems # empty list = may be booked
What should happen when a rule fails? Not a quiet correction. Code that overwrites the total with the sum of the lines assumes the total was misread, and on the supermarket receipt it was a line item. Your code should fail closed: the receipt goes to review with its problems listed next to the photo, and a person settles it in seconds. Some teams add one bounded repair attempt first: send the problems back, ask Claude to read the receipt again, and run the new reply through every gate from the start.
The gates between Claude and the ledger
stop_reason is end_turn6.3.5 Skepticism: confidence is not evidence
Now the taxi receipt, which no parser or arithmetic rule would ever catch. Why would Claude write a VAT number that is not there? A language model generates the most plausible continuation of what it has seen. When the output must contain a VAT number for a German taxi receipt, a well-formed German VAT number is a very plausible thing to write. Sunniva's first prompt asked for a VAT number and offered no way to say "none printed", so the model filled the box. An invented value arrives in the same format, with the same assured tone, as a value read off the paper.
Think of a fluent new hire who fills in every box on every form, because an empty box feels like failure. Their forms look perfect. Asking them to be more careful changes little; letting them leave a box empty, and checking the boxes that carry money against the records, changes a lot. Anthropic's guidance on reducing hallucinations (content that sounds right but is false or unsupported by the input) starts from the same place: explicitly give Claude permission to say it does not know. The nullable vat_number with its description does exactly that. The same guidance warns that such techniques reduce hallucinations without eliminating them, so critical information must still be validated.
So skepticism toward confident output is a design rule, not a mood: decide which facts carry real consequences, and check those against something other than the model. A confidence score you ask Claude for is generated the same way as the VAT number, and "are you sure?" gets an equally fluent answer. Running the same receipt several times and comparing is another technique the documentation suggests: disagreement can point to an invention, though agreement proves little. When the source is text rather than a photo, ask for the exact quote behind each claim and have your code confirm it appears in the source.
Match the check to the stakes. Obinna does not care whether a line reads "Espresso" or "Coffee", but a wrong VAT number costs the agency money and trouble with the tax office.
| Field | What rides on it | How the pipeline checks it |
|---|---|---|
| Line item descriptions | A label for the reader | Accepted as read |
| Date | The expense policy | Business rule: not in the future, inside the claim window |
| Totals | Money booked to the ledger | Arithmetic in code, within a cent |
| VAT number | The tax claim | Null allowed; must match the vendor's record or an official registry lookup, otherwise a person checks it |
Remember the principle rather than the rows: the higher the stakes, the more independent the check.
6.3.6 The exam traps
Every trap here trusts some property of the reply that does not prove it is safe to act on: that it was requested politely, that it parsed, that it sounded sure.
- ✗ Asking for "JSON only" in the prompt and parsing whatever comes back. ✓ Send a schema with JSON outputs, or use a strict tool when the data is an action, so the shape is constrained rather than requested.
- ✗ Treating a schema-valid reply as correct data. ✓ Check business rules in code. The schema guarantees the shape, with named exceptions, and never the arithmetic or the truth.
- ✗ Parsing before reading
stop_reason, or closing the brackets on a truncated reply. ✓ In an extraction call, onlyend_turnis a finished reply. Raisemax_tokensor split the job; a patched array silently loses data. - ✗ Defaulting a missing or unexpected value so the pipeline keeps moving. ✓ Fail closed and hold the item with its reason. A default of 0, "approve" or "resolved" turns a detected error into an action.
- ✗ Trusting a confidence score, an "are you sure?" follow-up or a bigger model to stop inventions. ✓ Allow "not present", and verify the facts that matter against a source you control.
- ✗ Pulling fields out of a friendly explanation with a regular expression. ✓ Keep the machine-readable payload in its own schema. If people need an explanation too, give it a field of its own.
Four reasons to trust a reply, one that counts
confidence: 0.98more generated text6.3.7 Put it together: build a receipt checker, then fool it
You now have every piece: an enforced schema, a parser that assumes nothing, business rules in code, verification for the facts that matter, and a review queue for the rest. The quickest way to make it stick is to build the pipeline on three receipts, then fool it on purpose and watch which gate catches what.
The same boundary appears across the rest of this guide. Debugging and error handling (4.1) covers what to do next with a held item: which failures deserve a retry, and how to tell a bug in your integration from a mistake in the output. AI application security (7.1) looks at the receipt itself as a threat, such as text in an uploaded photo that tries to give Claude instructions. Tool implementation (8.1) returns to strict schemas from the tool side, where a guaranteed shape still needs a sense check in your handler.
Key takeaways
- ✓ Output handling is the set of checks between Claude's reply and anything that acts on it, because a reply that parses can still be incomplete, inconsistent or invented.
- ✓ A prompt only requests a shape; JSON outputs (
output_config.format) constrain the reply to your schema, andstrict: truedoes the same for a tool's input. - ✓ The schema guarantee has named exceptions (refusals,
max_tokensstops, enum capitalisation), so readstop_reasonfirst and treat onlyend_turnas a finished extraction. - ✓ Parse defensively: find the text block by type, validate in one guarded step, and never patch truncated JSON or default a missing field.
- ✓ A valid shape is not valid data: check business rules such as totals in your code, and hold anything that fails for review with its reason.
- ✓ Confidence is not evidence: allow "not present", and verify high-stakes facts against your records or the source, never against the model's view of itself.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
18 CCDV-F questions on Domain 6, free
Every question in the bank is tagged to a domain, so you can drill 18 questions on Prompt and Context Engineering alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 6 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.