Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAO-F › Domain 2 › Lesson 2.1

CCAO-F · Domain 2 · 21% of the exam · Lesson 2.1 · 19 min read

Checking outputs for accuracy and completeness

Why a polished Claude output can be accurate yet incomplete, how to check coverage against a list you write first, and which claims to spot-check.

Written against objective 2.1 of the official CCAO-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.1.1 Why a finished-looking table can still let you down

You are a business analyst at your city's transport authority, and the authority is buying a new ticketing app. Three vendors have answered its request for proposals (RFP), the document that sets out the 15 requirements every bid must meet. The evaluation panel meets on Thursday. Each proposal runs to about 80 pages, so you give Claude the RFP and all three bids and ask for a comparison table the panel can score from.

The table that comes back is a pleasure to read. Requirements sit under clear themes, every vendor gets a short verdict per row, and a closing summary recommends Vendor B as the only bid that meets every requirement.

It also has three problems, and none of them shows. The table has 13 rows, not 15: accessibility and data retention, both set out in an annex at the back of the RFP, never made it in. And the support row credits Vendor B with round-the-clock help, when its proposal offers 07:00 to 19:00 on weekdays. The rows are grouped by theme rather than numbered, so nothing hints at a gap, and the wrong cell looks exactly like the right ones.

None of this is exotic. Anthropic's Help Center says plainly that Claude can give incorrect or misleading answers, that it should not be your single source of truth, and that high-stakes advice deserves careful scrutiny. The skill in this lesson is knowing that scrutiny means two separate checks. Accuracy asks whether what is on the page is correct. Completeness asks whether everything that should be on the page is actually there.

2.1.2 Is it right, and is it all there?

Here is the belief that catches people out: if everything in a draft is correct, the draft is fine. It isn't, because correctness only covers what the draft says. Anthropic's AI Fluency course calls the skill of judging output discernment, and it asks both whether there are factual errors and whether there are gaps in the analysis. A draft can pass the first test and fail the second by leaving out a requirement, a condition or an exception the reader needed.

Look at Vendor C's row on fare capping, the rule that stops charging a rider once they reach a daily limit. The cell says "Supported", and so does the proposal, which adds that capping works only for riders with a registered account, not for people tapping a bank card. The RFP asks for capping for every rider. Every word in the cell is true, and a panel reading it would still score Vendor C wrongly.

A witness in court swears to tell the truth AND the whole truth: two promises, not one, because someone can say only true things and still mislead by leaving out what matters. An output owes you both, and each check needs something different.

Accuracy Completeness
The question Is what is there correct? Is everything that should be there actually there?
A failure looks like A wrong number, name, date or condition, or a conclusion the evidence does not support A missing requirement, section, condition or exception
Visible on the page? Sometimes, if you know the subject No: nothing marks what is absent
Checked against The source A list made before reading: the brief, the requirements, the source's sections
In the vendor table Vendor B's support hours Accessibility, data retention, the capping condition

The row to memorise is the fourth: accuracy is checked against the source, completeness against a list you made first.

Polish makes the second check easy to skip. Anthropic's AI Fluency Index, a study of nearly 10,000 conversations on Claude.ai, found a pattern worth knowing. When Claude produced a finished piece of work, such as a document, code or an app, people were less likely to identify missing context, check facts or question Claude's reasoning. One possible explanation the researchers give is that if the work looks finished, people treat it as finished. Their advice fits your table: when something looks good, that is the moment to ask whether it is accurate, whether anything is missing and whether the reasoning holds up.

2.1.3 Omissions are silent, so write the list first

Why couldn't you have caught the missing rows by reading more carefully? Because reading shows you what is in the output, and a missing requirement isn't in it. However carefully you read those 13 tidy rows, you are checking the table against the table.

Think of a delivery arriving at the office. You don't check it against the packing slip in the box, because the slip lists only what was packed. You check it against your purchase order, written before anything shipped. Claude's output is the packing slip. Put precisely: completeness can only be checked against a list made independently of the output, before you read it, so the output cannot shape what you expect. Usually that list already exists: the brief you wrote, the requirements you were given, the sections of the source.

Two ways to check the vendor table for gaps

Reread the table

13 tidy rows
Each row looks rightchecked against itself
Two gaps unseen

Check against the RFP

List R1 to R15 firstannex included
Tick each against the table
R9 and R13 have no row
Rereading checks the table against itself, so missing rows stay invisible. A list taken from the RFP first makes every gap show up.

Here, the list is the RFP's 15 requirements, annex included, plus what the panel needs in every row: a verdict per vendor, any conditions, and a page reference. That is a checklist, or a rubric once you add what counts as a pass for each item. When Anthropic asked experienced Claude users how they check its work, the barrier they named most often was a lack of subject expertise. It is hard to judge an answer when you are not sure what good looks like. A checklist written in advance is what good looks like, on paper.

You can also put the checklist into the prompt, so gaps surface on their own. Look at the second and fifth lines: one row per requirement in RFP order, and an allowed answer for silence.

Compare the three attached proposals against the RFP's 15 requirements, listed below as R1 to R15. R9 (accessibility) and R13 (data retention) come from Annex C.
Give one row per requirement, in RFP order, with the requirement number in the first column. Keep every requirement as its own row, even when two seem related, so the panel can score each one.
For each vendor, give a verdict (Meets, Partly meets, Does not meet or Not addressed) and the proposal page it comes from.
If a vendor meets a requirement only under a condition, such as an account type, a contract length or an extra fee, put the condition in the cell.
Use "Not addressed" when a proposal says nothing about a requirement, because the panel scores silence as a gap.
After the table, list any requirement you could not assess, and why.
(R1 to R15 pasted from the RFP)

A skipped requirement now shows up as a missing number or an explicit "Not addressed", instead of as nothing. The prompt makes your check fast without replacing it, because any instruction can be followed imperfectly. Tick the returned table against your own list before you judge a single verdict.

2.1.4 Checking what is there

Once every row exists, the question is whether you can believe them. You can't reread three 80-page proposals, and you don't need to: the claims that carry weight are specific enough to check against the source in a minute or two.

Specific claims are also where errors hide, because a plausible wrong figure looks exactly like a right one. "Vendor B: 24/7 support" reads as confidently as "Vendor A: 24/7 support", and only one of them matches its proposal. The table lists the kinds worth checking.

Kind of claim In the vendor table Check it against
Numbers Each vendor's five-year cost The price schedule in each proposal
Names and dates The go-live date each vendor commits to The timeline page Claude cited
Quotes and commitments "24/7 support" for Vendor B Vendor B's support section
Conditions "Price includes hosting" for Vendor A The exact wording: for how long, and what hosting costs after that
Totals and comparisons "Vendor A is cheapest over five years" Your own sum of the figures
Conclusions "Only Vendor B meets every requirement" The corrected table: does it still hold?

Memorise the six kinds: whatever the output, these are the claims worth checking.

The last two rows go beyond single facts. A set of correct figures can still produce a wrong total or ranking, so add them up yourself. And a conclusion can stop following from the evidence once you fix one cell. Correct Vendor B's support hours and it no longer meets requirement 11, which asks for live support whenever the network runs, evenings and weekends included. The recommendation, the line the panel will act on, falls with that one cell.

The page numbers in your checklist prompt make this fast: each check means opening one page, not searching eighty. Anthropic's guide to reducing hallucinations (confident details that are not true) recommends the same habit: have Claude back each claim with a quote or source, so every claim can be traced. A reference speeds up the check without replacing it, because the page may not say what the cell claims.

2.1.5 Long outputs: check what matters most, sample the rest

The full table has 45 verdict cells, plus prices, conditions and a summary. Checking every one would take the rest of the week; checking none would hand an unchecked table to a panel about to award a multi-year contract. Spend your time where an error would do the most damage.

Start with the claims the decision rests on: the recommendation, every pass or fail on a mandatory requirement, and the prices. Then check a small sample of the remaining cells, picked at random rather than the ones that are quickest to confirm. If the sample turns up an error, widen the check to the whole row or column, since one error suggests others from the same reading.

Claude can help you decide where to look. Anthropic's Claude 101 course advises verifying key facts yourself for high-stakes work, and asking Claude to cite sources or say how confident it is. The AI Fluency Index suggests asking it outright what it is uncertain about. For the vendor table, that takes one follow-up message, and the first three questions do the work.

Before I check your table, help me decide where to look.
1. Which of the RFP's 15 requirements does the table leave out, or cover only in part?
2. Where a proposal was vague, what did you assume?
3. Which five cells are you least sure of, and why?
4. For each answer, give the proposal page I should open.

To see what the answer is worth, try it on the first, 13-row table, where you already know the three problems. The reply is useful: it notes that data retention sits in Annex C and never made it into the table, and it names Vendor C's pricing as the cell it is least sure of. But it says nothing about accessibility, and it doesn't flag Vendor B's support hours, the cell that was actually wrong.

A student who marks the exam answers they are unsure of helps the teacher, but a student who misread a question is sure of the wrong answer. Claude's account of its own work comes from the same reading that produced the table. It points to places worth checking, but it cannot vouch for the places it didn't mention.

Where your checking time goes

DECISION CLAIMSrecommendation, pass or fail, prices
CLAUDE'S FLAGSgaps, assumptions, least-sure cells
RANDOM SAMPLEa few other cells, not the easy ones
WIDEN IF WRONGthe whole row or column
SIGN OFFknowing what you checked
Check the claims the decision rests on, then the places Claude flags, then a random sample of the rest, and widen the check wherever something turns out wrong.

2.1.6 The exam traps

Nearly every trap in this objective lets the output check itself. Only the right answer checks it against something outside the output: the requirements, or the source.

  • ✗ Treating a polished, professional-looking output as finished. ✓ Check it against the requirements. Formatting is evidence of formatting, not of coverage.
  • ✗ Rereading the output to see whether anything is missing. ✓ Check against a list made before reading. An omission leaves nothing on the page to notice.
  • ✗ Signing off because every claim you checked was correct. ✓ Check completeness separately: accurate content that omits a requirement, condition or exception can still be unusable.
  • ✗ Asking Claude whether its output is complete and accurate, and taking yes for an answer. ✓ Ask what it left out and where it is least sure, then check those places and your own list. Its answer is a pointer, not proof.
  • ✗ Checking only the first few rows, or the ones quickest to confirm. ✓ Check the claims the decision rests on first, then sample the rest at random, and widen the check if the sample finds errors.
  • ✗ Throwing the output away, or switching models, after finding a gap. ✓ Fix the specific gaps and check the fixes. The failure was coverage, not the tool.

Four easy responses, one real check

Send itit looks professional
Reread itgaps leave no trace
Ask "is it complete?"a self-report
Reformat itchanges the look only
Check against the list and the sourcecoverage item by item, key claims page by page
Sending, rereading, asking Claude and reformatting all leave the gaps where they are. Only a check against something outside the output finds them.

2.1.7 Put it together: check a comparison before it reaches the panel

You now have the whole method. Before Thursday, you list R1 to R15, annex included, and rerun the comparison with one row per requirement. Accessibility and data retention appear, and Vendor C turns out never to address retention at all, a finding the first table hid. You check the decision claims against the cited pages, correct Vendor B's support hours and watch the recommendation fall. Then you check where Claude is least sure and sample the rest. The panel gets a table you can stand behind, and you know exactly what you checked.

The rest of Domain 2 builds on these two checks. Hallucinations, inconsistencies and bias (2.2) are particular ways accuracy fails, each with its own warning signs. Fact-checking and validation (2.3) takes the accuracy check beyond the documents you gave Claude, to outside sources and the question of which to trust. And deciding when human review is needed (2.4) sets how deep these checks go for what is at stake.

Key takeaways

  • ✓ Accuracy asks whether what is there is correct; completeness asks whether everything that should be there is there, and a draft can pass one while failing the other.
  • ✓ A polished, professional-looking output is evidence of formatting, not of accuracy or coverage.
  • ✓ Omissions leave no trace on the page, so check coverage against a list made before reading: the brief, the requirements or the source's sections.
  • ✓ Turn the brief into a checklist and ask for one entry per item with "not addressed" allowed, so a gap becomes visible instead of silent.
  • ✓ Spot-check the specific claims (numbers, names, dates, quotes, conditions) against the source, then check totals, comparisons and whether the conclusion still follows.
  • ✓ In a long output, check the claims the decision rests on first, sample the rest at random, and widen the check when the sample finds an error.
  • ✓ Claude's list of gaps, assumptions and least-sure items tells you where else to look; it never replaces the check.

Check your understanding

4 questions written for this lesson, then one from the CCAO-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

78 CCAO-F questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 78 questions on Output Evaluation and Validation alone, or sit the full 60-question timed simulator.

Open the CCAO-F question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources