Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 5 › Lesson 5.5

CCAR-F · Domain 5 · 15% of the exam · Lesson 5.5 · 21 min read

Human review workflows and confidence calibration

Why a high overall accuracy can hide a failing document type or field, and how calibrated confidence, routing and stratified sampling aim human review.

Written against task statement 5.5 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.5.1 When can you stop checking the machine's work?

Picture a new clerk in accounts payable, the team that pays suppliers, keying invoices into the finance system. For their first month a supervisor checks every invoice they type, and the error log looks excellent: 98 values in every 100 are right. Should the supervisor stop checking? That depends entirely on WHERE the other two are. If they're scattered across harmless fields, perhaps. If they're packed into the due dates on one supplier's smudged paper invoices, the company will pay that supplier late every month, and nobody will notice until the reminder letters arrive.

An extraction system built on Claude puts you in the supervisor's chair. The model reads each invoice and fills in a structured record: invoice number, dates, line items, total. A schema, the template that fixes which fields the record has and what type each one holds, can guarantee the record's shape. It can't tell you whether the values are right: a wrong due date looks exactly like a correct one. Nor can the model hand you its own error rate. It gives its best reading and, if you ask, a number for how sure it is, and that number is a claim, not a measurement.

So the real design question isn't whether the model is accurate. It's which extractions a person should look at, and how you'll know when that answer changes. A human review workflow answers it with evidence: accuracy measured segment by segment against known answers, confidence scores tested against those answers, and a steady sample of the work nobody reviews any more.

5.5.2 The average that hides a broken segment

Here is the situation the exam loves. A wholesale distributor runs its supplier invoices through a Claude extraction pipeline, where a tool called extract_invoice turns each invoice into a record of six fields, from invoice_number to due_date and total. Four reviewers check every one of the roughly 12,000 invoices a month before anything is paid. Against 2,000 recent invoices whose correct values the reviewers confirmed, the pipeline gets 98% of field values right. The finance director asks: can we stop reviewing most of them? (Every figure in this lesson is illustrative.)

Before anyone answers, take the 98% apart. An aggregate accuracy figure is one average over everything, weighted by volume: every invoice counts once, so the common kinds count most. Seven invoices in ten are clean, typed PDFs from large suppliers, which the model reads well, so they dominate the number. Split by document type, one supplier stands out. Kestrel Packaging still posts paper invoices, scanned on arrival; they are one invoice in twenty, and on those scans about one value in five is wrong.

Split by field, due_date is right only 93% of the time, with misses at every supplier. Many invoices print payment terms such as "Net 30" (pay within 30 days) instead of a date, or a date like 04/05/2026 that could mean April or May.

One number versus the same results split up

All invoices one figure

98% of values right
"Safe to cut review"the tempting conclusion

Split by document type

Typed PDFs, large suppliers70% of invoices, 99% right
Emailed PDFs, small suppliers25% of invoices, 98.6% right
Kestrel's scanned invoices5% of invoices, 81% right

and due_date is 93% right, across all suppliers

The 98% is mostly the large, easy segment. Split by document type, one supplier's scanned invoices get about one value in five wrong.

Weighted by volume, those three segments average out to 98%, and Kestrel's failure moves the headline by less than one point. Think of a river that is knee-deep on average. People still drown in it, because the average includes the wide shallows and says nothing about the channel in the middle. Cut review on the strength of the aggregate, and Kestrel's invoices and every doubtful due date flow straight into the payment run while the dashboard still says 98%.

So analyse accuracy by segment before you reduce human review. A segment is any slice of the work that might behave differently: a document type, a supplier or template, a field. Compute accuracy for every segment and for the combinations that matter, such as Kestrel's due dates. Automate only the segments that consistently clear the bar the business needs, and keep full review on the rest until a fix brings them up. A segment with a dozen examples proves nothing either way, so your known answers must include enough of the rare document types to measure them. Kestrel's 100 invoices in the set are enough to show the problem.

5.5.3 Confidence for every field, not every invoice

Segment analysis tells you which KINDS of invoice are risky, not which value on this morning's invoice is doubtful. For that you need a signal on every extraction: a field-level confidence score. It's a number from 0 to 1 next to each value, saying how sure the model is that the value matches the document. It isn't automatic. You add the field to the extract_invoice schema and explain it in the prompt, and the model fills it in along with the value.

Here's part of one record as your code receives it, in JSON (JavaScript Object Notation), the plain-text format programs use to pass data. Look at the due_date lines: a low score, and a note saying why.

{
  "invoice_number": { "value": "BW-1187",    "confidence": 0.98 },
  "invoice_date":   { "value": "2026-03-02", "confidence": 0.96 },
  "due_date":       { "value": "2026-05-04", "confidence": 0.58,
                      "note": "printed 04/05/2026: 4 May or 5 April?" },
  "total":          { "value": 4180.00,      "confidence": 0.97 },
  "conflict_detected": false
}

Why per field? Because the doubt lives in one value. One score for the whole record forces a bad choice: send the invoice to a person, who re-checks every sound value to find the doubtful date, or wave it through, date and all. Field-level scores let a reviewer check just the flagged value. They also let each field have its own threshold, the cut-off score below which a value goes to a person. That matters, because fields fail in different ways.

Tell the model to lower the score when a value is inferred rather than printed, when the print is unclear, or when the document disagrees with itself. Anthropic's guidance on reducing hallucinations makes a related point: give Claude explicit permission to say it isn't sure.

But what does 0.9 actually mean? As it comes, not much. Raw self-reported confidence is poorly calibrated: it isn't a measured probability, the model can be confident and wrong, and the hard cases are exactly where that happens.

5.5.4 Calibration: checking what the scores are worth

Here's the question that separates a design from a guess: when the model says 0.9, how often is it right? The only way to know is to compare its scores with answers you already have. That comparison is calibration: finding out what each score is actually worth.

Think of a weather forecaster. Collect every day she said "70% chance of rain" and count the days it rained. If it rained on about 70 in 100, her 70% means what it says, and she is well calibrated. If it rained on only 40, her 70% is overconfident, and you learn to read it as 40. Calibrating confidence scores is the same bookkeeping: group extracted values by the score the model gave, then measure how often each group was right.

The known answers come from a labelled validation set: real documents for which people have recorded the correct value of every field. "Labelled" means each document carries its correct answers, its labels. "Validation" means the set is used only for checking; if the same invoices also shaped the prompt and its examples, the check would flatter you. Think of the answer key to an exam paper, kept away from the students. The set should mirror production, Kestrel's scans and odd date formats included, with enough of each that the result isn't luck. The distributor already has one: the 2,000 invoices its reviewers confirmed.

The model said Actually right on the validation set Read it as
0.95 to 1.00 99.6% Meets a 99.5% bar: safe to auto-accept
0.90 to 0.94 86% Sounds safe, but wrong about one time in seven
0.80 to 0.89 74% Needs a person
below 0.80 51% Close to a coin toss

Memorise the method, not these illustrative numbers. Notice what the table does to the team's first instinct, a threshold of 0.9 picked by feel. Values scored 0.90 to 0.94 are wrong about one time in seven. The pipeline handles about 72,000 values a month (12,000 invoices, six fields each). If just 3% of them land in that band, a 0.9 threshold lets about 300 wrong values a month into the payment run.

The business sets the bar, say 99.5% right for anything auto-accepted, and the threshold is where the lowest band that clears it begins: 0.95 here. Build the table for each field, because the same score means different things on different fields; totals clear the bar at 0.95, due dates only at 0.98. And recalibrate whenever the prompt, the model or the document mix changes, because the scores' meaning belongs to the setup that produced them.

5.5.5 Routing: spending a small team's hours where the doubt is

With calibrated thresholds you can answer the finance director, and the answer isn't "stop reviewing". It's "review differently". Four people checking 12,000 invoices a month means 3,000 each, and effort spread that evenly lands mostly on values that were already right, 98 in every 100. Routing sends each extraction straight through or to a person, by three rules.

  1. Low confidence. Any field below its calibrated threshold goes to a reviewer, who sees that value highlighted next to the document instead of re-checking the whole record.
  2. Ambiguous or contradictory source. Some documents don't settle the answer themselves. The header says "due 30 September" while the payment slip says "due 15 September"; the printed total doesn't match the line items; a date reads two ways. The extraction sets conflict_detected, and the record goes to a person whatever the confidence.
  3. Segments that failed. Before you automate a segment, rerun the segment check on just the values that clear their thresholds, by document type and field. On typed and emailed invoices those values meet the 99.5% bar. On Kestrel's scans even the confident values fall short, because the model misreads smudged digits without knowing it. So Kestrel's invoices stay in full review until a fix, such as examples of their layout in the prompt, brings them over the bar.

Everything else is accepted automatically. The second rule is the one people get wrong. Why route a contradiction the model is confident about? Because confidence describes the reading, not the document. The model can be sure it read both dates correctly; what it can't know is which date the supplier meant, and settling that takes a phone call or a business decision.

How each extraction is routed

Extracteach value with its confidence
Comparecalibrated threshold, conflict flag
Reviewlow score or contradiction
Accepteverything else goes through
Samplea few accepted ones per group
Sample → Extract · sample findings feed calibration
Each field is compared with its calibrated threshold; doubtful fields and contradictory documents go to a person, and a sample of what is accepted is checked too.

In this example, about 2,000 of the 12,000 invoices a month now reach a person, all 600 of Kestrel's among them. That is 500 each for the four reviewers instead of 3,000, usually for one field rather than six. Each correction is also one more known answer for the next calibration. Notice what no rule does: none deletes or blanks a value because its score is low. Confidence decides who checks a value, never whether it is kept. That is prioritising limited reviewer capacity: attention follows the doubt, not the volume.

5.5.6 Stratified sampling: measuring what nobody reviews

Now the uncomfortable question: once most invoices go straight through, how do you know they're still right? The calibration was measured once, on last quarter's invoices. Since then suppliers have changed templates, new ones have arrived and someone has edited the prompt. Worse, the errors that matter most are the confident ones, because they sail past every threshold. Only a person looking at accepted records can catch a confident mistake.

So a person keeps looking, at a sample. A plain random sample, say 100 accepted invoices a month, is mostly typed PDFs from large suppliers, because most invoices are. The small groups, where problems tend to start, get a handful each or none.

The fix is to split before you sample. Picture a fruit inspector facing a delivery of 90 crates of apples and 3 crates of mangoes. A random handful from the whole load is nearly all apples, and the mangoes may never be tasted. So the inspector picks a few at random from every kind of crate.

That is stratified random sampling. You divide the accepted, high-confidence extractions into groups called strata (the word means layers) by what might make them fail: document type, supplier group, how new the supplier is. Each invoice belongs to exactly one group. Then you draw a random sample from EACH group, sized so its error rate can be measured. Picking at random within a group matters too, because it stops anyone choosing "the ones that look odd" and skewing the rate.

Group of accepted invoices (share) Plain random 100 a month Stratified sample a month
Typed PDFs, large suppliers (76%) about 76 40
Emailed PDFs, small suppliers (21%) about 21 40
Suppliers new this quarter (3%) about 3 30

Three invoices can't give an error rate: one mistake reads as 33%, none as perfect. Thirty can at least show a problem. Each sampled invoice is checked field by field, so you get field error rates too. Kestrel's scans need no sample, because all of them are reviewed.

The sample does two jobs. It keeps a current error rate for every group of work nobody otherwise sees. And it catches novel error patterns, mistakes of a kind the validation set never contained. In week seven, it turns up three invoices from a large supplier that moved its delivery date to where the due date used to be. The model copied it as the due date with 0.99 confidence, above even the strict 0.98 due-date threshold, so routing could never have caught it. The team routes that supplier to review, adds the invoices to the validation set, fixes the prompt and recalibrates.

5.5.7 The exam traps

Every trap here is one mistake in different clothes: trusting a number nobody has checked against known answers. The exam usually shows you the mistake and asks for the fix.

  • ✗ Cutting review because overall accuracy is high. ✓ Break accuracy down by document type and field first, and automate only the segments that meet the bar. The average is weighted toward the easy majority.
  • ✗ Routing on the model's raw confidence, with a threshold picked by feel. ✓ Calibrate on a labelled validation set and set each field's threshold from measured accuracy. Self-reported confidence can be high and wrong.
  • ✗ One confidence score for a whole document. ✓ A score per field, so reviewers check the doubtful value and each field gets its own threshold.
  • ✗ Reviewing a random 1% of everything as quality control. ✓ Stratified random sampling of high-confidence extractions, so small groups get enough checks to be measured.
  • ✗ Validating once before launch, then trusting the thresholds forever. ✓ Keep sampling. New layouts and suppliers create error patterns the validation set never saw, often with high confidence.
  • ✗ Auto-accepting a contradictory document because the model is confident. ✓ Route ambiguous or contradictory sources to a person whatever the score; confidence describes the reading, not the document.

Wrong ways to decide what gets reviewed

Overall accuracyhides the weak segment
Raw model confidencenever checked against answers
A random 1% of everythingtoo few from small groups
One check before launchmisses new error patterns
Segment checks, calibrated routing, stratified samplingall measured on labelled data
Each wrong answer leaves a weak segment, a bad score or a new error unseen. The right one measures by segment, calibrates, and keeps sampling.

5.5.8 Put it together: build a review router and break it

You now have every piece, from the average that hides a broken segment to the sample that keeps measuring what nobody reviews. Building a miniature version makes each failure visible in a way no table can.

The same discipline closes out this domain. Provenance and multi-source synthesis (5.6) moves from invoices to research reports, and the moves will feel familiar. You keep a record of which source supports each claim, annotate conflicts instead of quietly picking a side, and report gaps in coverage rather than letting a confident summary hide them.

Key takeaways

  • ✓ An aggregate accuracy figure is weighted toward the common, easy documents, so it can hide a document type or field that fails badly.
  • ✓ Before reducing human review, measure accuracy by document type and field against known answers, and automate only the segments whose high-confidence values meet the bar.
  • ✓ Have the model return a confidence score with every field so review can target the one doubtful value, but treat the raw score as a self-report, not a probability.
  • ✓ Calibrate on a labelled validation set: measure how often each confidence band was actually right, and set each field's review threshold from that measured accuracy.
  • ✓ Route fields below their calibrated threshold, and documents that are ambiguous or contradictory, to people, so a small team spends its time where the doubt is.
  • ✓ Keep stratified random sampling of auto-accepted, high-confidence extractions running, so every group has a current error rate and new error patterns surface early.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

54 CCAR-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 54 questions on Context Management & Reliability alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources