Home › Study guides › CCAR-P › Domain 2 › Lesson 2.3
CCAR-P · Domain 2 · 13% of the exam · Lesson 2.3 · 22 min read
Zero-shot, few-shot and chain-of-thought: choosing the technique per decision
When clear instructions are enough, when examples must teach your conventions, when reasoning earns its tokens, and how an eval set settles the choice.
Written against objective 2.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
2.3.1 Why one prompting technique cannot serve every decision
Stallwick is an online marketplace for second-hand and handmade goods. About 60,000 new listings arrive every day, and each one needs up to three decisions before it goes live. First, which of 38 shop categories it belongs in. Second, whether it carries one of Stallwick's own restriction codes: R1 (the courier checks the buyer's age at delivery), R2 (the seller must upload a document first) or R3 (not allowed at all). Third, for about one listing in a hundred, a judgement against the listing policy: a 1950s folding knife described as "a collector's piece, handy for self-defence", or turmeric capsules "clinically proven to ease arthritis".
Idris, who leads trust and safety, built the prototype. One prompt handled all three decisions, with a line asking Claude to reason step by step in its reply and twelve past listings pasted in as examples. Categories came back right, but slowly and at a cost nobody had budgeted. Every kitchen knife was flagged R1, borderline rulings changed from one run to the next, and some came back refused. Idris's instinct was to move everything to the most capable model. Philippa, the architect brought in to take the system to production, asked a different question first: what does each decision need from the prompt?
There are three classic answers. Zero-shot prompting gives Claude instructions and context but no worked examples. Few-shot prompting (Anthropic's docs also call it multishot) adds a handful of example inputs paired with the right outputs. Chain-of-thought asks Claude to reason before it answers, and on current models a built-in feature called thinking does this natively. Each technique buys something and costs something: examples add input tokens to every call, reasoning adds output tokens and seconds. The architect's job is to choose per decision, then prove the choice on an eval set, a batch of real cases with known right answers.
Three decisions, three techniques
Category
Restriction code
Borderline ruling
2.3.2 Zero-shot: clear instructions and the baseline
Here is the belief that trips people up: zero-shot sounds like "the prompt with nothing in it", so it must be the weak option. In fact a good zero-shot prompt is full. It says what to produce and why, defines every allowed output, and says what to do when an input fits nothing. The only thing it leaves out is EXAMPLES. Anthropic's prompting guide calls this being clear and direct, with a golden rule: if a colleague with minimal context would be confused by the prompt, Claude will be too.
Zero-shot is enough when three things hold: the task rests on general knowledge, the outputs can be defined in words, and the eval shows the prompt meets its bar. Stallwick's categories meet all three, since any shopper can tell a teapot from a tent peg. In the excerpt, look at the first line, which gives the reason behind the task, and at the last two, which cover the silent cases where zero-shot prompts usually fail.
You assign each new Stallwick listing to exactly one shop category. Buyers browse by category, so a wrong category hides the item from the people looking for it.
<categories>
Kitchen & Dining: cookware, tableware, kitchen tools and small kitchen appliances
Garden & Outdoor: plants, garden tools, outdoor furniture, barbecues
(36 more categories, one line each)
</categories>
Choose by what the item is, not by words in the title: a "garden gnome mug" belongs in Kitchen & Dining.
If two categories fit, choose the one a buyer would search first. If none fits, answer "Other".
Zero-shot is also the baseline. It spends the fewest input tokens, has no example set to keep current, and cannot teach a pattern you did not intend. Every other technique must beat it on the eval to earn a place. Philippa ran the category prompt on Claude Haiku 4.5, thinking off, against 400 moderator-labelled listings. It scored 97.6% against a bar of 97%, so it stayed zero-shot.
2.3.3 Few-shot: teaching the house's own conventions
The restriction codes are a different problem. A zero-shot prompt with careful definitions scored 78%, and the errors were not random. At Stallwick a non-locking pocket knife under 7.5 cm is clear, while any locking blade is R1; a sunscreen with an SPF claim is R2, while a plain moisturiser is clear. Those lines were drawn case by case in years of moderators' rulings, and no general knowledge contains them, so a bigger model would not know them either. The prompt was short of information, not of intelligence.
Few-shot prompting is how a new moderator really learns: a senior colleague walks them through five tricky past cases. Examples teach three things at once: the output format, what your labels mean in practice, and where the boundary between two labels lies, which definitions struggle to pin down. Anthropic's guide calls examples one of the most reliable ways to steer format, tone and structure. It recommends three to five, each relevant (close to real inputs), diverse (covering edge cases, varied enough that Claude picks up no unintended pattern) and structured (in <example> tags inside <examples>, apart from the instructions).
Look at the pairs in Philippa's set: each boundary appears from both sides, and the restricted examples share no surface feature.
<examples>
<example><listing>Folding pocket knife, non-locking, 8.5 cm blade, beech handle</listing><code>R1</code><why>Over 7.5 cm, so R1 even though it does not lock.</why></example>
<example><listing>Pocket multi-tool, 5.8 cm non-locking blade, scissors, nail file</listing><code>CLEAR</code><why>Non-locking and under 7.5 cm.</why></example>
<example><listing>SPF 50 mineral sunscreen, sealed, 100 ml</listing><code>R2</code><why>SPF claims need the seller's product safety certificate.</why></example>
<example><listing>Handmade shea moisturiser, no SPF or medical claims</listing><code>CLEAR</code><why>Cosmetics without claims are clear.</why></example>
<example><listing>Replica hand grenade, inert, display piece</listing><code>R3</code><why>Replica explosives are never allowed.</why></example>
</examples>
Examples also mislead, because Claude learns from everything in them, including what you never meant to teach. Most of Idris's restricted examples were knives, so the model learned "knife means restricted": it flagged a butter-knife set and missed a can of lighter refill gas, which Stallwick sells with an age check. His set was also mostly restricted, which taught that restriction is normal. And it was copied unreviewed, old mistakes included.
Two example sets for the same codes
Twelve pasted past listings
flags butter knives, misses lighter gas
Five curated examples
96.5% on the same eval
Philippa's five lifted the restriction eval from 78% to 96.5%, above its 96% bar, for about 250 extra input tokens a call. She versions them like code and keeps them out of the eval set, which would otherwise grade the model on answers it had been shown.
2.3.4 Chain-of-thought: reasoning before the ruling
Now the borderline queue, about 600 listings a day. Take the 1950s knife. Its locking blade makes it R1 under clause 4.2, and clause 4.6 lets collectors' pieces through with an age check. But the seller wrote "handy for self-defence", and clause 4.5 bans any item marketed as a weapon: R3. The right ruling needs steps: identify the object, find every clause that applies, notice the marketing language, and apply the house rule that the strictest outcome wins. A model that answers in a single pass has to compress all of that into its first word, like a moderator stamping the form before reading it through.
Chain-of-thought comes in three forms. The first two are prompts. The third is thinking, where the model reasons in separate thinking content blocks that arrive before its text answer. On current models thinking is adaptive: Claude decides per request whether and how deeply to think.
| Form | How you ask for it | Where the reasoning lands |
|---|---|---|
| Basic | "Think it through step by step, then answer" in the prompt | In the reply text, mixed with the answer |
| Structured | Reason inside <thinking> tags, then answer inside <answer> tags |
In the reply text, in separate tags your code can split |
| Native thinking | Adaptive thinking, steered by the effort parameter; on older models such as Claude Haiku 4.5, extended thinking with a token budget |
In thinking blocks, separate from the text answer |
Memorise the rule behind the rows: native thinking is the default on the newest models, and the prompt forms are the fallback where thinking is off, structured tags preferred because your code can split them. Thinking is always on for Claude Opus 5.5 and the Fable models. On Claude Sonnet 5.5 it is on by default, and its lowest setting, between_tools, only drops the thinking before the answer, so a request without tools gets no reasoning at all. On Claude Haiku 4.5, thinking is off unless you enable extended thinking.
Effort is the main dial for adaptive thinking. At low, Claude skips thinking on simple tasks; at high, it thinks on most requests that benefit. The default is high on most models but medium on Claude Opus 5.5, so set it explicitly. The docs favour general guidance over scripted steps, so Philippa's prompt lists the clauses a ruling must weigh, not a script of thoughts.
One old habit now backfires. On Claude Opus 5.5, Claude Sonnet 5.5, Claude Fable 5 and Claude Fable 5.1, a prompt asking Claude to write its reasoning into the reply can be declined. The response ends with stop_reason: "refusal" and the category reasoning_extraction. Asking Claude to think before answering is fine, because with thinking on that only steers the thinking. Anthropic's advice is to remove instructions that stood in for thinking and read the reasoning from thinking blocks, so Idris's "reason step by step in your reply" line, the cause of his refusals, goes.
Reasoning is not free. Thinking tokens are billed as output tokens even when the thinking text is hidden, they count toward max_tokens, and they add response time. Each response reports the spend in usage.output_tokens_details.thinking_tokens. Reasoning earns its cost where the answer depends on working out: rules with exceptions, conflicting facts, calculations. On a lookup or a house convention it usually adds seconds, not accuracy. On 50 borderline cases, Claude Opus 5.5 at high effort agreed with senior moderators on 46, against 37 at low.
2.3.5 Keep the reasoning out of the answer your application uses
Once Claude reasons, what does your application actually read? The prototype read the whole reply, a paragraph of deliberation followed by a verdict, and its parser broke whenever the wording shifted. Worse, one early build pasted that text into the rejection notice, telling sellers exactly which phrases tripped the weapons clause: a free lesson in evading moderation.
The fix is a clean split, and thinking already provides it: reasoning in thinking blocks, the answer in a text block. Add structured outputs, a JSON schema sent in output_config.format, and the text block can hold only JSON matching your schema. The schema constrains Claude's direct output, not its thinking, so Claude reasons freely and still answers in your exact shape. The flip side: all the working out now happens in thinking, so an effort level low enough to skip thinking costs accuracy. With thinking off, the <answer> tag plays the same role: your code extracts that tag and nothing else.
In this sketch, look at the clause field (a justification by citation, not a transcript), the lookup by block type, and the log that keeps thinking internal.
RULING = {"type": "object", "additionalProperties": False,
"required": ["ruling", "clause"],
"properties": {"ruling": {"type": "string", "enum": ["CLEAR", "R1", "R2", "R3", "REFER"]},
"clause": {"type": "string"}}} # cite the clause, no reasoning transcript
def rule(listing_id: str, listing_text: str) -> dict:
response = client.messages.create(
model="claude-opus-5-5", max_tokens=16000, system=POLICY_PROMPT,
thinking={"type": "adaptive", "display": "summarized"}, # return a readable summary
output_config={"effort": "high", "format": {"type": "json_schema", "schema": RULING}},
messages=[{"role": "user", "content": f"<listing>{listing_text}</listing>"}],
)
if response.stop_reason != "end_turn": # refused or cut off: a moderator decides
return {"ruling": "REFER", "clause": ""}
details = response.usage.output_tokens_details
log_internal(listing_id, [b.thinking for b in response.content if b.type == "thinking"],
details.thinking_tokens if details else 0) # reviewers only, never the seller
return json.loads(next(b.text for b in response.content if b.type == "text")) # by type
On Claude Opus 5.5 the thinking text is hidden unless you ask for it; display: "summarized" returns a readable summary, which goes to an internal log for moderators and prompt engineers. It is never the raw chain of thought, so it is a debugging aid, not the record of the decision: that is the ruling and its cited clause. And because adaptive thinking can skip thinking on an easy request, code that assumes a fixed first block breaks.
Where the reasoning goes
<listing> tagsR3, clause 4.52.3.6 Choose, combine, and prove it on an eval set
With three techniques on the table, the tempting design is all three, everywhere, to be safe. Resist it. At 60,000 listings a day, a technique that adds nothing still bills every call.
| Technique | When it wins | What it costs |
|---|---|---|
| Zero-shot | The task rests on general knowledge, outputs can be defined in words, the eval meets the bar | Fewest tokens, least upkeep; misses conventions the prompt never states |
| Few-shot | Your own labels, formats or boundaries that definitions do not pin down | Input tokens on every call; an example set to curate and version; can teach unintended patterns |
| Chain-of-thought | Multi-step judgements: rules with exceptions, conflicting facts, calculations | Output tokens and latency; the answer must be kept apart from the reasoning |
| Combined | A judgement that also follows house conventions | The costs add up, so each part must show its gain on the eval |
Memorise the middle column; the last is how you justify the choice to a finance lead. To decide what to add, read the baseline's errors. Errors clustered on your conventions or on format call for examples, and errors on cases needing several steps call for reasoning. Errors because a fact is absent, such as a policy clause never put in the prompt, call for context, not a technique. Techniques also combine: Stallwick's borderline route uses the clauses as instructions, three past rulings as examples of how clauses interact, and thinking at high effort.
The eval set makes each choice defensible. Anthropic's guidance is to make evals task-specific (the real distribution plus the edge cases), to automate grading where you can, and to prefer many automatically graded cases over a few hand-graded ones. Stallwick grades codes and categories by exact match in code, and borderline rulings by exact match plus a rubric check on the clause. Every run records accuracy, p95 latency and cost per listing, thinking tokens included.
Start zero-shot, step up on evidence
Philippa's decision record fits on one page, each line with its evidence.
DR-07 Listing moderation: prompting technique per decision. Owner: Philippa (architect). Approved: Idris (trust and safety).
Category, 60,000 a day: zero-shot on Claude Haiku 4.5, thinking off. 97.6% on 400 labelled listings, bar 97%. Five examples added 0.2 points for about 250 input tokens a call: rejected.
Restriction code, 60,000 a day: five curated examples on Claude Haiku 4.5. Zero-shot 78%, few-shot 96.5%, bar 96%. Example set versioned; no example is an eval case.
Borderline ruling, about 600 a day: policy clauses, three example rulings, adaptive thinking at high effort on Claude Opus 5.5, JSON ruling with cited clause. 46 of 50 agree with senior moderators, bar 45. Thinking logged internally, never shown to sellers.
Review: re-run all three evals on any change to a prompt, an example set or a model.
2.3.7 The exam traps
Every trap here applies a technique to the wrong failure, or without evidence.
- ✗ Adding examples and a reasoning instruction to every prompt "to be safe". ✓ Start zero-shot and add a technique only where the eval shows a gap it closes.
- ✗ Asking for step-by-step reasoning to fix wrong house labels or an inconsistent format. ✓ Add examples. Reasoning cannot recover a convention the prompt never states; examples show it.
- ✗ Pasting in a large batch of past decisions as examples. ✓ Curate three to five relevant, diverse examples with no dominant label, and keep them out of the eval set. Raw history carries its skew and its old mistakes.
- ✗ Moving to a bigger model when the model is missing your conventions. ✓ Supply the missing information with examples. No model knows what your moderators decided last year.
- ✗ Prompting the newest models to write out their reasoning in the reply. ✓ Use thinking and effort, and read summarized thinking where you need it. Such prompts can be declined as
reasoning_extraction; tagged manual reasoning is for models with thinking off. - ✗ Letting users, parsers or audit records consume the model's deliberation. ✓ Your code reads a structured answer by block type or tag, and people see a short justification that cites the rule.
2.3.8 Put it together: pick and prove a technique per decision
You now have every piece: a zero-shot baseline, examples for conventions, reasoning for multi-step judgement, the answer kept apart from the reasoning, and an eval set deciding each step. Now measure the techniques on your own small data, and watch a bad example set undo the gain.
The rest of the domain builds on these choices. Context and token management (2.4) weighs examples and thinking against everything else the context window must hold, and prompt reuse (2.5) caches the stable prefix where a fixed example block belongs. Evaluation datasets (4.2) and A/B testing (4.3) turn your small eval set into a standing discipline.
Key takeaways
- ✓ Choose the prompting technique per decision, not per application, and start from the simplest one that could work.
- ✓ A zero-shot prompt carries clear instructions, context and a definition of every output; it is enough when the task rests on general knowledge and the eval meets the bar.
- ✓ Few-shot examples teach your own labels, formats and boundaries: use three to five relevant, diverse examples in tags, with no dominant label or shared surface feature.
- ✓ Chain-of-thought pays on multi-step judgements; its native form is thinking, billed as output tokens and steered by effort on the newest models, with tagged manual reasoning as the fallback where thinking is off.
- ✓ Keep the reasoning apart from the answer your application reads: thinking blocks or tags for the reasoning, a structured answer with a short cited justification for your code and your users.
- ✓ Prove each choice on an eval set kept apart from the examples, compare accuracy, latency and cost per call, and adopt the cheapest configuration that meets the bar.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
24 CCAR-P questions on Domain 2, free
Every question in the bank is tagged to a domain, so you can drill 24 questions on Claude Models, Prompting & Context Engineering alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 2 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.