Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 4 › Lesson 4.4

CCAR-F · Domain 4 · 20% of the exam · Lesson 4.4 · 21 min read

Validation, retry and feedback loops for extraction quality

Retry an extraction with the validation error attached, know when a retry cannot help, and add calculated_total, conflict_detected and detected_pattern.

Written against task statement 4.4 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

4.4.1 Why a well-formed extraction can still be wrong

Picture a temp worker keying supplier invoices into an accounts system, with a supervisor who checks each batch before anything is paid. When an invoice is wrong, the supervisor can hand it back with "wrong, redo it", and the temp will probably retype it the same way, because nothing says where the slip is. Or the supervisor can attach a note: "the lines add up to 1,240.00, but you typed 1,420.00 as the total". That takes ten seconds to fix. And if the note says "payment terms missing" when the terms are printed only in a supply agreement nobody put in the folder, no amount of redoing will produce them.

An extraction system built on Claude meets the same three situations. Your code sends Claude the invoice together with a tool, a named form Claude can fill in. The tool's JSON schema is the layout of that form: each field, the kind of value it holds (text, number, date) and which fields are required. JSON (JavaScript Object Notation) is the plain-text format programs use to pass data to each other. Claude answers by calling the tool, and the input of that call is your structured record.

The schema makes sure the record has the right SHAPE. It cannot make sure the record is RIGHT. A total can be misread, a date can land in the wrong field, and nothing in a single model call checks your business rules. The fix is a validation-retry loop. Your code runs every extraction through a validator, a set of checks you write. When a check fails, your code sends the model the document, its own answer and the exact error, so the model can correct itself. Two supporting ideas sit around that loop. You design the output so it exposes its own inconsistencies, and you record enough about each result to learn where the system goes wrong.

4.4.2 Two kinds of wrong: syntax and semantics

Here is the distinction the rest of this lesson depends on, and the exam leans on it. Some outputs are broken as data. Others are perfectly good data that says something false, and the two need different cures.

A schema syntax error means the output is not usable as data. Either the JSON itself is broken, with a missing bracket or a sentence of prose wrapped around it. Or it breaks the schema: a total sent as the text "1,240.00" where the schema wants a number, or a required field left out. Tool use with a JSON schema is the cure. The answer arrives as the input of a tool call shaped by your schema, so this whole class of failure goes away, and with it the need to retry for it.

A semantic error is well-formed data with the wrong meaning. Take invoice INV-2291 from an office supplier: ten chairs for 890.00, five lamps for 210.00 and delivery at 140.00, with a printed total of 1,240.00. The extraction comes back valid against the schema, but with stated_total: 1420.00 (two digits swapped) and the due date copied into invoice_date. Every type is right; the values are wrong. In the temp's terms, every box on the form is neatly filled in, and one of them holds the wrong number.

Kind of error On invoice INV-2291 What removes or catches it
Syntax: malformed output a missing bracket, text around the JSON Tool use: the answer is a tool call's input
Syntax: breaks the schema total as the text "1,240.00", invoice_number missing The tool's JSON schema
Semantic: values disagree lines sum to 1,240.00, stated_total says 1,420.00 Your validator, then a retry with the error
Semantic: wrong field the due date in invoice_date Your validator ("invoice date must come before due date"), then a retry

Memorise the split: syntax belongs to tool use and the schema, semantics to your validator. In Python, teams often build the validator with Pydantic, a widely used free library (a package of ready-made code) for checking data. You describe the record once, field by field: stated_total is a number, invoice_date is a date, payment_terms may be empty. Pydantic then checks each extraction against that description and reports every failure as a readable message that names the field. You can add your own rules to the same description, such as "the line items must add up to the total", which a schema's type rules cannot express.

4.4.3 Retry with the error attached

Your validator has caught INV-2291: stated_total is 1420.00 but the line items add up to 1240.00. What do you send next? The tempting answer is the same request again, perhaps with "please be careful with numbers" added. That tends to earn the same kind of answer, because the model still has no idea which part was wrong.

The technique that works is retry with error feedback: a follow-up request that carries three things.

  1. The ORIGINAL document. The API keeps no memory between requests, so the model can only re-read what you send again.
  2. The FAILED extraction. Seeing its own answer lets the model fix the broken field and keep the rest, instead of starting over and making a new, different mistake.
  3. The SPECIFIC error. "stated_total is 1420.00 but line_items sum to 1240.00" names the field and the contradiction. "Validation failed" names nothing.

This is the supervisor's note from the opening, turned into code. Anthropic's prompting guide calls the general pattern self-correction: generate an answer, review it against criteria, refine it, each step a separate call. Here the reviewer is your validator, so the criteria are exact.

The validation-retry loop

EXTRACTtool call returns the record
VALIDATEschema checks plus business rules
RETRYdocument + failed output + error
ACCEPTevery check passes
RETRY → EXTRACT · on a validation error, up to a small cap
Every extraction goes through the validator. A failure sends the document, the failed output and the exact error back to the model, and your code validates the corrected answer again.

In code, the three numbered lines are the whole technique; the cap and the final flag keep the loop honest.

errors = validate(extraction)          # Pydantic or JSON Schema, plus your business rules
attempts = 0
while errors and attempts < 2:         # a small cap: retries are for fixable errors
    attempts += 1
    followup = (
        f"<invoice>\n{invoice_text}\n</invoice>\n"                  # 1. the ORIGINAL document
        f"<extraction>\n{json.dumps(extraction)}\n</extraction>\n"  # 2. the FAILED extraction
        "Validation failed:\n- " + "\n- ".join(errors) + "\n"      # 3. the SPECIFIC errors
        "Re-read the invoice, fix only these problems, and call extract_invoice again."
    )
    extraction = call_extract_invoice(followup)   # same tool, same schema
    errors = validate(extraction)
if errors:
    flag_for_review(extraction, errors)           # still failing: a person looks at it

If you continue the original conversation instead of building a fresh request, the same three pieces are already present. The document sits in the first message, and the failed extraction is Claude's own tool call. You answer that call with a tool_result block (the message that carries a tool's outcome back to Claude), put the error text inside and mark it is_error: true. Either way, write the error as an instruction a colleague could act on.

4.4.4 When a retry cannot help

Now the harder question: is this error worth retrying at all? A retry is the model reading the same document a second time with a pointer to the problem. So it can only succeed when the right answer is IN that document.

That covers a lot of ground. A date written "15/09/2026" where your validator wants "2026-09-15" is a format mismatch: the value is on the page, just in the wrong form. A delivery charge placed inside another line item, or a due date in the invoice date field, is a structural error: the value is there, in the wrong place. Swapped digits are a misreading. In every case the specific error sends the model back to text that contains the answer.

Absence is different. INV-2291's payment terms line reads "as set out in the attached supply agreement", and the agreement was never sent. The validator reports payment_terms as empty, you retry, and the model re-reads the same invoice and finds the same sentence. No note can make the temp read a page that is not in the folder. Worse, if the field is required, repeated demands push the model toward a plausible invention such as "Net 30", which passes validation and is wrong.

Will a retry fix it?

Retry fixes it

Format mismatch"15/09/2026" for an ISO date
Structural errora value in the wrong field
Misread value1,420.00 for 1,240.00

retry with the specific error

Retry cannot fix it

Value in an attachmentthe supply agreement, not sent
Page missing from the scan
Never written down at all

stop: null, flag, supply the source

A retry succeeds when the answer is in the document and was misread, misformatted or misplaced. It cannot succeed when the information is not there.

For absence, the right move is to stop retrying. Let the field be null (empty on purpose), flag the record with the reason, and either fetch the missing document and extract again with it included, or route the record to a person. Anthropic's advice on reducing hallucinations points the same way: give Claude explicit permission to say the information is not there.

A quick test decides whether to retry: could a careful human holding only this document fix the error? If yes, retry. If no, retrying just spends calls. Over time, log which kinds of error a retry actually fixes. That record tells your code which errors to retry and which to send straight to review.

4.4.5 Fields that make the extraction check itself

Some problems stay invisible because the record has nowhere to show them. Suppose the schema has a single total field and an invoice's printed total does not match its own lines. Should the model copy the printed figure or write the correct sum? Whichever it picks, the other number is lost, and nothing in the record says the invoice disagreed with itself. The fix is to design self-correction fields: pairs of values that should agree, so that a disagreement becomes visible to your code.

The classic pair is stated_total, the total exactly as printed, and calculated_total, the sum of the line items as extracted. To fill in the second, the model has to add up what it read, so the record carries its own cross-check. Your code compares the two, and can recompute the sum from the line items to check the model's arithmetic too. That comparison is how your validator caught INV-2291: stated_total 1,420.00 against calculated_total 1,240.00.

A mismatch has two possible causes, and they need different handling. If the model misread a number, as it did with INV-2291's total, a retry with the error fixes it. If the invoice itself does not add up, the extraction is correct and the mismatch is a real finding: the supplier's mistake. Never ask the model to "fix" the stated total to match. You would hide exactly the problem that accounts payable, the team that pays suppliers, most wants to see.

The same thinking covers sources that contradict themselves. INV-2291's header says "Due 30 September", but the payment slip at the bottom says "Due 15 September". Left alone, the model picks one date and nobody learns that a choice was made. A conflict_detected field, a boolean (true or false), gives the model a permitted way to say "this document disagrees with itself". It usually comes with a short note naming what disagrees.

Your validator then treats the two cases differently. A total mismatch with conflict_detected false is an error to retry. A record with conflict_detected true skips the retry and goes to a person, note attached. An honest flag reaches a reviewer; a silent mismatch never gets through.

Field What it holds What it catches
stated_total The total exactly as printed Nothing alone; it is the baseline for the comparison
calculated_total The sum of the extracted line items A misread or dropped line, or an invoice that does not add up
conflict_detected True when two parts of the document disagree Contradictory source data the model would otherwise resolve silently

4.4.6 Feedback loops: learning from dismissed findings

Everything so far fixes one record at a time. A feedback loop improves the system, and the exam places it in the other scenario of this domain: an automated code reviewer running in continuous integration (CI). CI is the automatic checking that runs whenever a developer proposes a change to the code, a proposal called a pull request. The reviewer posts findings on each pull request, comments such as "this error is caught and then ignored". Developers accept some findings and dismiss others.

Those dismissals are your best evidence of false positives: findings that flag a problem that is not really there. But the evidence only helps if you know what triggered each finding. "28% of findings dismissed this month" tells you the reviewer is noisy, not where.

The fix is a detected_pattern field on every finding: a short label for the code construct that triggered it, meaning the kind of code involved. Keep the labels a fixed set, an enum (a list of allowed values) in the schema, so the same construct always gets the same label and the counts mean something.

{
  "file": "billing/export.py",
  "line": 88,
  "severity": "minor",
  "message": "Exception is caught and ignored, so failures will be silent.",
  "detected_pattern": "broad_exception_catch"
}

Think of a restaurant that only counts plates sent back to the kitchen. It knows diners are unhappy; it does not know the fish is the problem. Note which dish each plate held and the answer jumps out. Group a month of dismissals by detected_pattern and you might find that most come from broad_exception_catch. Meanwhile sql_string_concat (a database query glued together from pieces of text, a real security risk) is almost never dismissed. Now you can rewrite the criteria for that one pattern, or switch it off while you do, and leave the findings developers trust untouched.

The dismissal feedback loop

TAGeach finding carries detected_pattern
DECIDEthe developer accepts or dismisses
ANALYSEdismissal rate per pattern
ADJUSTrewrite criteria for noisy patterns
ADJUST → TAG · measure again after each change
Tagging each finding with the pattern that produced it turns a pile of dismissals into a per-pattern false-positive rate you can act on.

4.4.7 The exam traps

Each trap below pulls a lever built for a different kind of failure.

  • ✗ Retrying with the identical request, or with "that was wrong, try again". ✓ Send the original document, the failed extraction and the specific validation error. Without the error, the model does not know what to change.
  • ✗ Raising the retry count when a value is missing from the document. ✓ Recognise absence: make the field nullable, flag the record, and supply the missing source or route it to a person. More attempts cannot find what is not there and invite a fabricated value.
  • ✗ Tightening the schema, with more required fields or stricter types, to fix totals that don't add up. ✓ Tool use and a schema remove syntax errors only. Semantic errors need validation code and a retry with the error.
  • ✗ Extracting only the printed total, or letting the model silently pick between conflicting values. ✓ Extract calculated_total beside stated_total and add conflict_detected, so discrepancies reach your code and your reviewers.
  • ✗ Tracking only the overall dismissal rate, or telling the whole reviewer to "be more conservative". ✓ Add detected_pattern to each finding and analyse dismissals by pattern, so you fix the categories that cause false positives.

4.4.8 Put it together: build a self-checking invoice extractor

You now have every piece. A small build makes the gap between a fixable error and a hopeless one concrete.

The rest of Domain 4 scales this up. Batch processing (4.5) runs the same extractor over thousands of invoices at half price, and resubmits only the documents that failed, identified by their custom_id, the label you attach to each request. Multi-instance review (4.6) adds a second, independent Claude instance to check what the first produced. And human review workflows (5.5) are where the records flagged here end up: stated and calculated totals that disagree, conflict_detected set, information that was never in the document.

Key takeaways

  • ✓ Tool use with a JSON schema eliminates schema syntax errors; semantic errors, such as values that don't sum or a value in the wrong field, still need your own validation.
  • ✓ A useful retry is a follow-up request with the original document, the failed extraction and the specific validation error.
  • ✓ Retries succeed on format mismatches and structural errors, because the right value is in the document.
  • ✓ Retries cannot succeed when the information is absent, for example in an attachment that was not provided; make the field nullable, flag the record and supply the source.
  • ✓ Self-correction fields expose discrepancies: calculated_total beside stated_total, and conflict_detected for sources that contradict themselves.
  • ✓ A detected_pattern field on each review finding lets you analyse dismissals by pattern and fix false positives where they come from.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

72 CCAR-F questions on Domain 4, free

Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 4 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources