Home › Study guides › CCAR-F › Domain 4 › Lesson 4.3
CCAR-F · Domain 4 · 20% of the exam · Lesson 4.3 · 20 min read
Structured output with tool use and JSON schemas
Why an extraction tool with a JSON schema beats asking for JSON, what tool_choice any and forced guarantee, and schema design that stops invented values.
Written against task statement 4.3 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.3.1 Why "reply in JSON" is not good enough
Picture the accounts-payable desk of a mid-sized company. Hundreds of supplier invoices arrive every week as PDFs and scans, each laid out differently, and someone types the supplier, invoice number, amounts and due date into the payment system. You want Claude to do the reading and a program to do the typing. The program is the fussy half: it needs every invoice in exactly the same shape, every single time.
The first thing everyone tries is to ask for it: "Extract this invoice and reply in JSON." JSON (JavaScript Object Notation) is the text format programs use to pass records around: named fields inside curly braces. The request works most of the time, and "most of the time" is the problem. Now and then the reply opens with "Here is the extracted data:", wraps the JSON in extra formatting, drops a comma or leaves out a field. The program that reads it fails, and at a few hundred invoices a week, one failure in fifty is several stuck invoices.
Think of the difference between a blank sheet and a printed form. Ask a temp to "write down the invoice details" and every temp hands back a slightly different page. Hand them a form with labelled boxes, tick boxes for the category and a box marked "not stated", and every page comes back in the same shape.
Tool use lets you give Claude that form. Normally a tool is an action your program offers Claude, such as looking up an order, and Claude requests it by filling in the inputs it needs. For extraction you define a tool whose inputs are exactly the record you want, Claude "calls" it by filling in the boxes, and your code reads the completed form.
4.3.2 An extraction tool is a form, not a function
Here is the part that surprises people: the tool you define for extraction never runs. There is no function called extract_invoice anywhere in your code, and for a one-shot extraction there is no result to send back either. The tool exists only so that Claude's answer has a contract.
A tool definition has three core parts: a name, a description and an input_schema. The input schema is a JSON Schema, a precise description of which fields an object has, what type each one is, and which ones must be present. For an ordinary tool, the schema describes the arguments a real function needs. For extraction you turn that around. The schema describes the record you want, so the "arguments" Claude writes ARE the extracted data. Claude reads the invoice and replies with a tool_use block, one of the separate pieces a reply is made of. Its input field holds the supplier, invoice number, dates, line items and total, already arranged the way the schema says.
Why is this more reliable than a polite request for JSON? Because the record never travels as prose you have to cut up. It arrives as a structured object in its own block. Add "strict": true to the tool definition and the API (application programming interface, the service your code calls to reach Claude) goes further. It constrains Claude's output while Claude writes it, so the input is guaranteed to match your schema. That rules out broken JSON, wrong types and missing required fields.
One invoice through the extractor
extract_invoice, call requiredtool_use blockinput is your JSONThe whole exchange fits in a dozen lines. Look at three of them: tools offers the form, tool_choice makes filling it in compulsory (the next section explains how), and block.input is the extracted record.
response = client.messages.create(
model=MODEL, max_tokens=2048,
system=EXTRACTION_RULES, # format rules (see below)
tools=[extract_invoice], # name, description, input_schema, strict
tool_choice={"type": "tool", "name": "extract_invoice"}, # the call is compulsory
messages=[{"role": "user", "content": invoice_text}],
)
block = next(b for b in response.content if b.type == "tool_use")
invoice = block.input # THIS is the data, shaped by the schema
check_invoice(invoice) # your own checks on the values
send_to_accounts_payable(invoice)
4.3.3 tool_choice: can Claude skip the form?
Offering a tool is not the same as getting a tool call. By default Claude decides for itself whether a tool is needed, and on an odd document it may decide it isn't. Given a smudged scan, it might reply in plain text: "This appears to be a delivery note rather than an invoice." That is a sensible sentence and a useless result: your code was waiting for a tool_use block that never came. The setting that closes this gap is tool_choice, a field on each request.
Stay with the paper form for a moment. auto says "fill in a form if you think one applies, or write me a note instead". any says "fill in one of these forms; you pick which". A forced tool says "fill in THIS form". In the API:
tool_choice |
What Claude must do | When to use it in extraction |
|---|---|---|
{"type": "auto"} |
Decide: call a tool or answer in text. The default when tools are provided. | When a text answer is acceptable, which in a pipeline it rarely is |
{"type": "any"} |
Call one of the provided tools; Claude chooses which. | Several extraction schemas exist and the document type is unknown |
{"type": "tool", "name": "extract_metadata"} |
Call that named tool. | One particular extraction must run before anything else |
{"type": "none"} |
Call no tool at all. | Not in extraction; it switches tools off |
Memorise the first three rows; the exam centres on them. none is there so you recognise it.
any fits the moment the accounts-payable inbox stops being tidy. Alongside invoices it now receives credit notes and receipts, so you define three tools: extract_invoice, extract_credit_note and extract_receipt. any guarantees that Claude calls one of the three and leaves the choice to Claude, who has read the document. Forcing extract_invoice here would push every credit note into the invoice form. On models that accept any, combine it with strict: true and you get both guarantees: a tool is called, and its input matches the schema.
Forcing a named tool is about order. Suppose the pipeline also enriches each document with tools of its own: it looks up the supplier in the vendor records and matches the purchase order before payment. Those steps need the supplier name, invoice number and purchase order (PO) reference first, so you add a small extract_metadata tool that captures just those. The first request sends tool_choice: {"type": "tool", "name": "extract_metadata"}. Your code reads the metadata and sends back a short tool result, and the follow-up requests return to auto so Claude can choose the enrichment steps. Forcing applies to one request. Leave it forced on every request and nothing but extract_metadata can ever run.
4.3.4 Valid is not the same as correct
A few weeks after switching to strict tool use, the parse errors are gone, not rare but gone. The accounts-payable team still finds bad records. One invoice has line items of 400.00, 590.00 and 250.00, which add up to 1,240.00, while the total field says 1,420.00. Another has the invoice date sitting in due_date, so the payment run thinks the bill is already overdue. Both records are perfect JSON that matches the schema in every respect.
A spell-checker makes the point well. It will happily pass "the meeting is on Tuesday" when the meeting is on Thursday, because every word is spelled correctly. A schema works the same way. It constrains structure: which fields exist, their types, which are required and which enum values are allowed. It knows nothing about how values relate to each other (do the line items add up to the total?) or what they mean (which of these dates is the due date?). Errors of structure are syntax errors; errors of meaning are semantic errors, and no schema can prevent them.
What a strict schema stops, and what it lets through
Syntax errors prevented
Semantic errors not prevented
due_datecaught only by checks you write
So the schema is the first layer, not the last. After reading block.input, your code runs its own checks: add up the line items and compare them to the total, confirm the due date is not before the invoice date. A record that fails never reaches the payment system; your code retries it with the error attached or sends it to a person. Clear field descriptions help as well: "due_date: the date payment is due, often after 'Due' or 'Pay by'; not the invoice date" gives Claude a reason to put each value in the right box.
4.3.5 Designing the form: nullable fields and escape hatches
The next problem is the most dangerous, because it looks like success. About a third of the invoices carry no purchase order number, yet every record in the payment system has one. The schema declared po_number as a required string, and strict mode guarantees exactly that: the field is present, and it is a string. On an invoice with no PO, Claude still has to write something, so it writes the next best thing: a reference number from the page, or a plausible "PO-10482". The schema demanded an answer the document could not give.
The fix is to make "not there" a legal answer, and there are two ways to do it. Leave the field out of required, which makes it optional, or keep it required but allow null as its value, which makes it nullable. Nullable is often the cleaner choice for a pipeline, because the field is always present and null plainly means "not on this document". Pair it with one line in the prompt, "use null for anything the document does not show", and Claude reports the gap instead of filling it.
Categories need the same care. An enum restricts a field to a fixed list of values, say goods, services and freight. That keeps the category column tidy until an invoice for customs duty arrives and gets squeezed into freight. A paper form's tick boxes end with "Other (please specify)", and your enum should too. Two extra values do the job:
otherplus a detail string. Claude can tell what the item is; it just isn't on your list. It picksotherand writes "customs duty" incategory_detail. The detail strings show you which categories to add later, so the list grows without breaking the schema.unclear. The document doesn't let anyone tell: "Q3 support pack" could be goods or services.unclearlets Claude say so, and your pipeline routes that record to a person instead of accepting a confident guess.
In the tool definition, look at three lines: po_number, which may be null; the category enum ending in other and unclear; and category_detail, the free-text line beside other.
{
"name": "extract_invoice",
"description": "Record the fields of one supplier invoice. Use null for anything the document does not show.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"total": {"type": "number"},
"po_number": {"type": ["string", "null"]},
"category": {"type": "string", "enum": ["goods", "services", "freight", "other", "unclear"]},
"category_detail": {"type": ["string", "null"]}
},
"required": ["invoice_number", "total", "po_number", "category", "category_detail"],
"additionalProperties": false
}
}
4.3.6 Formats: the rules that belong in the prompt
One more inconsistency slips past even a well-designed schema. Suppliers write dates as "3rd March 2026", "03/04/26" and "2026-03-04", and amounts as "€1,234.50" or "1.234,50 EUR". If invoice_date is a string, every one of those is a valid string. Two invoices from the same supplier can come back with their dates in two shapes, and the matching step downstream quietly fails.
The schema decides the shape of the record, and it can even require a date to look like a date. It cannot tell Claude how to read the messy original: whether "03/04/26" is the 3rd of April or the 4th of March, or which mark in "1.234,50" is the decimal point. Those conventions belong in the prompt, next to the schema, as explicit normalisation rules: how to read and write a date, an amount, a currency. It also helps to repeat the key rule in the field's own description, where Claude sees it at the moment it fills that box.
| On the invoice | Rule in the prompt | What the field receives |
|---|---|---|
| "3rd March 2026" | Dates as year-month-day, 2026-03-03 |
"2026-03-03" |
| "1.234,50 EUR" | Amounts as plain numbers with a decimal point; currency as a separate three-letter code | 1234.5 and "EUR" |
| "Net 30" | Payment terms as a number of days | 30 |
| "03/04/26" from a UK supplier | Read numeric dates by the supplier's country convention | "2026-04-03" |
Remember the pairing rather than the specific rules: the schema fixes which boxes exist, and the prompt fixes how to write inside them.
4.3.7 The exam traps
Every trap in this task statement comes from expecting one mechanism to do another's job: the prompt to guarantee structure, the schema to guarantee truth, or tool_choice to guarantee the right tool.
- ✗ Asking for JSON in the prompt and parsing the reply text. ✓ Define an extraction tool with an
input_schema(andstrict: true) and read thetool_useinput. A request for JSON is followed most of the time; a schema is enforced every time. - ✗ Leaving
tool_choiceonautoin a pipeline that needs structured output. ✓ Useanywhen several schemas exist and the document type is unknown, or force a named tool when one extraction must run first. Underauto, Claude may answer in text. - ✗ Using
anyto make one particular extraction run first. ✓ Force it with{"type": "tool", "name": "extract_metadata"}, then return toauto.anyguarantees some tool, not that tool. - ✗ Believing a strict schema fixes everything. ✓ It eliminates syntax errors only. Totals that don't add up and values in the wrong field need checks in your code, and inconsistent source formats need normalisation rules in the prompt.
- ✗ Marking every field required so the record looks complete. ✓ Make fields nullable when the document may not contain them. A required non-null field invites fabrication.
- ✗ A closed category enum. ✓ Add
"other"plus a detail string so the list can grow, and"unclear"so ambiguous documents are flagged instead of guessed.
4.3.8 Put it together: build an invoice extractor and break it
You now have every piece, from the extraction tool as a form to format rules in the prompt. The quickest way to make them stick is to build a small extractor and watch each safeguard fail when you take it away.
The rest of the domain builds on this extractor. Validation and retry loops (4.4) take over where the schema stops. They send the failed check back with the document so Claude can correct itself, and they recognise when a retry cannot help because the information is not there. Batch processing (4.5) runs the same extraction over thousands of invoices at lower cost, and records flagged unclear feed the human review workflows of Domain 5 (5.5).
Key takeaways
- ✓ Asking for JSON in a prompt works most of the time; the reliable route is an extraction tool whose
input_schemais the record you want, read from thetool_useblock'sinput. - ✓ The extraction tool never runs; with
strict: true, the API enforces its schema while Claude writes, so the output is schema-compliant and free of JSON syntax errors. - ✓
tool_choice:automay answer in text,anymust call one of the tools (several schemas, unknown document type), and{"type": "tool", "name": "extract_metadata"}must call that tool (one extraction before enrichment). - ✓ A schema prevents syntax errors, not semantic ones: line items that don't sum to the total and values in the wrong field need checks in your code.
- ✓ Make fields nullable when documents may not contain them; a required non-null field pushes the model to fabricate a value.
- ✓ Give category enums
"other"plus a detail string for extensibility and"unclear"for ambiguous cases. - ✓ Put format normalisation rules in the prompt alongside the schema: the schema fixes the fields, the rules fix how values are written.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
72 CCAR-F questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.