Home › Study guides › CCAR-F › Domain 4 › Lesson 4.2
CCAR-F · Domain 4 · 20% of the exam · Lesson 4.2 · 20 min read
Few-shot prompting for consistent, well-judged output
Two to four targeted examples fix what detailed instructions cannot: consistent review findings, fewer false positives, extraction without blanks or guesses.
Written against task statement 4.2 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.2.1 Why a page of rules still gives a different answer each time
Picture a new colleague whose job is to log customer complaints in a shared spreadsheet. You write them a careful page of rules: date, product, what went wrong, how serious it is. By Friday every row follows the rules and no two rows look alike: "urgent" in one, "high" in the next, a product name here, a product code there. Then you fill in three rows yourself, including an awkward complaint with a note on why you rated it low. From Monday their rows look like yours, and they rate a kind of complaint you never mentioned the way you would have.
A model working from instructions alone is that colleague on day one. Every instruction is words describing an output, and words leave room. "Give a severity" does not say which scale. "Include the location" does not say whether that means a file, a function or a line. The model fills each gap with a reasonable guess, and reasonable guesses differ from one run to the next. Adding more rules narrows the gaps without closing them, because each new rule is more words with gaps of its own.
The remedy is few-shot prompting (also called multishot prompting): you put a few worked examples into the prompt, each an input paired with exactly the output you want. One finished example answers dozens of questions you never thought to write down. Anthropic's docs describe examples as one of the most reliable ways to steer Claude's format, tone and structure, improving both accuracy and consistency. The best examples go further than a shape: they show a decision and the reason for it, which is what lets the model handle cases none of your examples covered.
4.2.2 Show the format instead of describing it
Continuous integration (CI) is the automated checking that runs whenever a developer proposes a change to shared code. The proposed change is called a pull request. Picture a team that runs Claude Code in its CI pipeline: on every pull request, Claude reviews the change and posts its findings for the developer.
The review prompt is detailed: for each finding, the location, the issue, a severity of low, medium or high, and a suggested fix. Yet the findings come back in a different shape each run. One is a paragraph, the next a bullet list with no line number; one severity reads "moderate"; one fix says "consider refactoring", which nobody can act on. The script that turns findings into pull request comments fails on half of them.
The team's first instinct is another paragraph of rules, and it helps for a while before the drift returns. The missing piece was never the list of fields. What was missing is a picture of a good finding: how precise a location is, how short an issue statement is, how concrete a fix must be. An example carries exactly that. That is why few-shot examples are the most effective technique for consistently formatted, actionable output once detailed instructions have proven not to be enough. Anthropic's consistency guide agrees: an example of the desired output works better than abstract instructions.
Here is one such example as it might sit in the review prompt. Notice the <example> tags, which Anthropic's docs recommend so the model can tell examples apart from instructions, and the suggested_fix line, which shows how concrete a fix has to be.
<examples>
<example>
<code>discount_per_item = discount / item_count</code>
<finding>
location: billing/cart.py, line 88
issue: item_count is 0 for an empty cart, so this line crashes checkout.
severity: high
suggested_fix: return the subtotal unchanged when item_count is 0,
before the division.
</finding>
</example>
... two more examples in the same shape ...
</examples>
With two or three examples like this, the output settles. A JSON schema can force the four fields to exist; an example shows what a good value in each one looks like.
Same instructions, with and without examples
Instructions only
same prompt, three shapes
Plus three examples
every finding matches the examples
4.2.3 Examples for the cases that could go either way
Format is the easy half. The harder half is judgement: what should the model do with a case the instructions do not clearly settle? Instructions are written for the typical case. The trouble lives at the edges, where two reasonable readings of one rule point in opposite directions.
Take two edge cases from the review pipeline. First, a pull request adds a login function that comes with tests. Code often splits into alternative paths, called branches: one for a valid login token, another for an expired one. None of the tests ever takes the expired-token branch. The instruction says "flag missing tests", the model reasons that the function is tested, and it stays silent. That is a branch-level test coverage gap, and whether it counts is exactly the kind of call an instruction leaves open.
Second, the same missing check appears in three files. The review agent has two tools for posting: post_inline_comment attaches a note to one line, and post_summary_comment posts one note on the whole pull request. Both are defensible here. That is tool selection for an ambiguous request: one request that plausibly fits two tools.
The skill is writing 2 to 4 targeted examples, each built around one genuinely ambiguous situation. Each example shows three things: the situation, the action chosen, and why the plausible alternative was rejected. The third part does the heavy lifting. "Report a coverage gap" is a verdict. "Report it, because coverage counts per branch and a tested function can still hide an untested path" is a rule the model can carry to the next case. For the three files, the example chooses one post_summary_comment listing all three locations, because three identical inline comments read as noise.
How one example teaches a judgement call
Why only a handful? Each example should earn its place by covering a different kind of hard case; eight examples of the obvious case make the prompt longer without making it wiser. For tool selection, examples also sit on top of clear tool descriptions, which Anthropic's tool docs call by far the most important factor in how well Claude uses tools. Examples settle the requests that stay ambiguous when every description is good. They cannot repair a description that never says what the tool is for.
4.2.4 Teaching judgement, not a lookup list
Here is the failure that costs a review system its credibility: false positives, findings that flag code which is actually fine. The tempting fix is to show the model more examples of real issues. The better fix is to show it pairs: an acceptable pattern next to a genuine issue that looks almost the same, with the reasoning that separates them.
In the review pipeline, the classic pair is an ignored error. When an operation fails, code can catch the failure and carry on as if nothing happened. In one place, the code deletes a temporary file and ignores the error if the file is already gone, with a comment saying so. In another, the code calls the payment service and ignores any error, so a failed charge looks exactly like a successful one. Same surface pattern, opposite verdicts. The example states why: ignoring an error is fine only when nothing depends on the operation having succeeded. With that pair in the prompt, the model stops flagging the harmless cases, which is where most of the false positives came from.
The reasoning matters even more for cases you never showed. Suppose a later pull request converts prices between currencies and quietly returns 0 when the exchange-rate service fails. No example mentioned currency, and the code looks nothing like the ignored-error examples. A model that learned the RULE sees the same shape underneath, a hidden failure in code whose result someone relies on, and flags it. A model given only verdicts learned something narrower, closer to "flag ignored errors, except the temp-file one". It keeps flagging harmless cases and walks straight past the currency fallback.
Think of training a football referee. Show them only which tackles were fouls and they memorise tackles. Tell them why each one was a foul (the player went for the leg, not the ball) and they can call a tackle they have never seen. Precisely: examples with reasoning let the model generalise its judgement to novel patterns, while answer-only examples teach it to match the specific cases you listed.
What the model learns from each kind of example
Answer-only examples
Examples with reasoning
One caution: Claude pays close attention to every detail of an example, including details you never meant to teach. If every flagged example comes from payment code, the model may learn "be strict with payment files". Anthropic's advice is to keep examples diverse, so vary the incidental details until the only thing your examples share is the rule.
4.2.5 Extraction: examples that prevent blanks and guesses
The review pipeline has a second job, and it puts the same technique to work on a different kind of task. Many pull requests link a short report that justifies the change, such as a benchmark: a timed test of how fast the new code runs. For each linked report, the pipeline extracts a JSON record for the team's dashboard: the sources the report cites, its sample size (how many runs the numbers rest on) and its measurements. It validates each record against a JSON schema before passing it on. This is structured data extraction: turning unstructured documents into fixed fields.
Extraction fails in two opposite directions. Either a required field comes back empty although the value is in the report, or the model invents a value for a field the report states only loosely or not at all. An invented value like that is a hallucination: output that looks confident but has no basis in the source.
Empty fields usually come from layout, not absence. The prompt says "extract the cited sources", and on reports with a numbered bibliography at the end that works. Reports that only cite inline, "(Okafor, 2021)" in the middle of a sentence, come back with an empty list, because nothing showed the model that inline references count. Likewise, reports with a Methodology heading yield a sample size, while reports that mention it in passing ("each figure is the average of 200 runs") leave a required field empty. One example of each layout, with the correct extraction, shows the model where the value hides. The fix is showing the variety, not rewording the rule.
Informal measurements invite the opposite failure. A report says the new version "now takes roughly a fifth of a second, give or take". Without guidance, the model may write a precise-looking "187 milliseconds" that nobody measured, or give up and leave the field blank. An example shows the handling you want: keep the value as stated, mark it approximate, quote the source words. Include one example where a field is genuinely absent and the right output is null, too. Anthropic's hallucination guide recommends giving Claude explicit permission to say it does not know; an example of an honest null is that permission, shown rather than told.
| Situation in the report | Without an example | What the example must show |
|---|---|---|
| Inline citations, no bibliography | An empty citations list |
Sources taken from "(Author, year)" in the text |
| Run count mentioned in the results, no Methodology heading | A required sample_size comes back empty |
The value pulled from a narrative sentence |
| An informal measurement, "roughly a fifth of a second" | An invented precise number, or a blank | The value as stated, marked approximate, with the source quote |
| The value is genuinely absent | A plausible invented value | null, because the report does not contain it |
Memorise the pattern, not the rows. Whether the documents are benchmark reports, research papers or invoices, you want one example for each layout they actually use, and one that shows what an honest blank looks like.
4.2.6 The exam traps
Every trap below pulls the wrong lever. Examples are the right lever when instructions are already detailed and the output still varies, or when hard cases need judgement. They are the wrong one when something else is missing.
- ✗ Adding another paragraph of rules, or "be consistent", when detailed instructions already produce inconsistent output. ✓ Add 2 to 4 examples in the exact target format. More words leave more gaps; an example closes them.
- ✗ Examples that show only the verdict. ✓ Show the reasoning, including why the plausible alternative was rejected. Verdicts alone teach surface matching; reasoning teaches judgement that generalises.
- ✗ Many examples of the obvious case, such as eight routine findings each routed to
post_inline_comment. ✓ A few targeted examples of the genuinely ambiguous cases. Obvious examples add length, not judgement. - ✗ Using examples to fix minimal tool descriptions. ✓ Rewrite the descriptions first, because they are what the model selects tools by. Examples then settle the requests that remain ambiguous.
- ✗ Relying on examples for a step that must happen every time, such as a check before a payment. ✓ Enforce it in code. Examples are guidance, and guidance is followed most of the time, not always.
- ✗ Fixing empty required fields on unusual layouts by making the field optional, retrying the same prompt, or ordering the model to always fill it. ✓ Add an example of the layout where the value hides. The value is in the document; the model did not know where to look, and a bare order to fill the field invites an invented value.
4.2.7 Put it together: turn a drifting prompt into a consistent one
You now have every piece. Examples fix the format, targeted examples with reasoning settle the ambiguous cases, contrasting pairs cut false positives and generalise, and layout examples keep extraction from leaving blanks or inventing values. The quickest way to believe it is to watch a prompt change when you add examples, and slide back when you remove the reasoning.
The rest of Domain 4 builds on these examples rather than replacing them. Structured output with tool use and JSON schemas (4.3) guarantees that every finding and every extraction arrives as valid JSON. Nullable fields in that schema give an honest blank a place to live, while examples still decide what good values look like. Validation and retry loops (4.4) catch what slips through, and a detected_pattern field on each finding shows which dismissed patterns deserve the next contrasting example. Before you push thousands of reports through the Message Batches API (4.5), you refine the prompt, examples included, on a sample set.
Key takeaways
- ✓ When detailed instructions still give inconsistent output, few-shot examples are the most effective fix: show the output instead of describing it.
- ✓ Examples in the exact target format (location, issue, severity, suggested fix) make every finding consistent and actionable.
- ✓ For ambiguous cases, such as an untested branch or a request that fits two tools, 2 to 4 targeted examples each show the situation, the choice and why the plausible alternative was rejected.
- ✓ Pairing an acceptable pattern with a similar genuine issue, reasoning included, cuts false positives and lets the model generalise to patterns it was never shown.
- ✓ In extraction, an example for each document layout stops required fields coming back empty, and examples of informal measurements and true absences reduce invented values.
- ✓ Examples refine judgement on top of clear tool descriptions and code-enforced guarantees; they do not replace either.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
72 CCAR-F questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.