Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 5 › Lesson 5.1

CCDV-F · Domain 5 · 16.8% of the exam · Lesson 5.1 · 22 min read

LLM fundamentals: tokens, sampling, thinking and examples

How tokens, context windows, sampling and next-token generation shape Claude's replies, and when to reach for thinking, effort, fast mode or examples.

Written against skill 5.1 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.1.1 Why a working summariser still surprises its developer

Dalia is the only developer at a four-person startup that turns meeting recordings into notes. A user uploads a transcript, and the app sends it to Claude with instructions: summarise the decisions, the action items and who owns each one. Users can then ask follow-up questions about the meeting. One Monday, Rohan, who runs the product, forwards three complaints.

First, their tester summarised one transcript twice and got two different summaries: the same decisions, different sentences, action items in another order. Dalia's regression test compares every summary with a saved copy, so it now fails about half the time. Second, summaries of the longest meetings stop in the middle of a sentence. Third, a user asked "Given the budget cut and the two deadlines, what should we do first?" and got three lines that restated the meeting instead of a plan.

It is tempting to treat these as three bugs and hunt through the code for each one. Resist that. The code does exactly what she wrote. All three come from how a large language model (LLM) works. It reads text as small pieces called tokens, inside a window of fixed size. It writes its answer one token at a time, picking each from a set of probabilities. A few settings on each request then decide how much room the reply gets, how much Claude thinks before answering and how fast it writes. Once you see the mechanism, each complaint has a fix of its own.

5.1.2 How Claude writes: one token at a time

Here is the belief behind most of the confusion: that Claude composes the whole summary in its head and then types it out. It does not. It writes the way you might speak without notes, choosing each next piece only after everything before it has been said.

A token is often a short word, sometimes part of a longer word or a punctuation mark. Everything is measured in tokens: what fits, what you pay and how long a reply takes. On current models, a million tokens hold roughly 555,000 English words. How text splits depends on the model's tokenizer, which changes between generations. Models from Claude Opus 4.7 on turn the same text into roughly 30 percent more tokens than older ones such as Claude Haiku 4.5. So never reuse a count; measure with the free token counting endpoint, for the model you will actually call.

Generation itself is a loop called next-token generation. Claude reads your prompt plus every token it has written so far, works out how likely each possible next token is, picks one and appends it. Then the loop runs again, until Claude ends its turn or the reply reaches the output limit you set.

Next-token generation

READthe prompt plus every token written so far
SCOREhow likely each possible next token is
PICKone token, sampled from those odds
APPENDadd it to the reply
STOPend of turn, or the max_tokens cap
APPEND → READ · repeat for every token of the reply
Each token is chosen after rereading everything before it and is then appended, so the reply grows one step at a time until Claude ends its turn or reaches your cap.

Three consequences follow, one for each complaint. A 1,200-token summary is 1,200 trips round the loop, so long replies take longer and a cap can stop one mid-sentence. A written token is never revised, so one different early pick sends the rest of the reply down a different path. And unless Claude reasons before it answers, the first sentence of a plan is committed before the rest has been thought through.

5.1.3 The context window: why long meetings get cut off

Dalia's first guess about the truncated summaries was that the transcript is too long for the model. Let's test that guess, because two different limits are easy to confuse here.

The context window is all the text the model can reference while generating a reply, including the reply itself. Think of it as working memory, not as the knowledge the model was trained on. Everything in the request counts: the system prompt, the tool definitions and every message. That includes earlier questions and answers, because the API remembers nothing between calls, so the app resends the conversation with each follow-up. Claude's output counts too, thinking included. The current Opus and Sonnet models have a window of 1M (one million) tokens; Claude Haiku 4.5 has 200k.

The second limit is much smaller, and it is yours: max_tokens, the most tokens Claude may write in this reply. Any thinking Claude does is billed as output and spends from that same budget. Dalia's app sends max_tokens=1024 to Claude Sonnet 5.5, a model that thinks by default. Her longest transcript, a two-hour planning meeting, is about 36,000 tokens, a small fraction of the 1M-token window. The window was never the problem. Every response carries a stop_reason field that says why Claude stopped, and the truncated ones all say "max_tokens": a long meeting needs a long summary, and any thinking spends from the same 1,024 tokens.

What shares one context window

Input what you send

System prompt and tool definitions
The transcript
Earlier questions and answersresent on every turn

Output what Claude writes

Thinkingwhen Claude thinks
The summary

capped by max_tokens, far below the window

Everything you send and everything Claude writes must fit in the window, while max_tokens separately caps what Claude writes, thinking included.
What you see The limit you hit What to do
stop_reason: "max_tokens" and text that ends mid-sentence Your max_tokens cap on the reply, which thinking shares Raise max_tokens or ask for a shorter output; never use the cut-off reply as complete
A 400 invalid_request_error: "prompt is too long" The input alone is larger than the context window Send less: trim or split the input
stop_reason: "model_context_window_exceeded" Input plus the reply so far filled the window Treat the reply as truncated and send less input

Memorise the first two rows; recognise the third. And fitting is not the same as helping: accuracy and recall degrade as the token count grows, which Anthropic calls context rot, so put in the window what the task needs, not everything you have.

5.1.4 Sampling: why two runs give two summaries

Now the first complaint: the same transcript and the same prompt, two different summaries. Nothing in Dalia's code changed between the runs, so where does the difference come from?

It comes from the PICK step of the loop. Claude does not always take the single most likely next token; it samples, choosing among the likely candidates according to their odds. Ask a colleague to retell a meeting twice and you get the same facts in different sentences. Sampling is that, done deliberately, one token at a time, and it is why LLM output is non-deterministic: the same input can give a different output on every run. Because each token depends on the ones before it, one different early pick ("The team agreed" instead of "Decisions:") reshapes everything that follows.

Three request parameters have traditionally shaped the pick. Temperature sets how adventurous it is: lower values stick to the most probable tokens, higher values allow more variety. top_p and top_k trim the candidate list first.

Parameter What it does On Claude Opus 4.7 and later models
temperature Randomness of each pick, from 0.0 to 1.0; the default is 1.0 Only the default is accepted; any other value returns a 400 error
top_p Samples only from the smallest set of likely tokens whose probabilities add up to this value Deprecated the same way: only values of 0.99 and above pass, so leave it out
top_k Samples only from the K most likely tokens Any value returns a 400 error

Memorise what temperature does; recognise the other two. Two facts limit all three knobs. First, no setting makes output identical: Anthropic's API reference warns that even a temperature of 0.0 is not fully deterministic. Second, as the last column shows, newer models such as Claude Sonnet 5.5 refuse them outright, and Anthropic recommends steering those models through the prompt instead. Older models such as Claude Haiku 4.5 still accept temperature, where a low value makes wording less varied, never fixed.

So the fix for Dalia is not a knob. It is to design for variation: pin the format, show examples of it (more on that below) and test the properties every good summary must have, not its exact words. Look at the two all(...) checks, which test format and facts on every run, and at the commented-out line, which is the test she had.

REQUIRED_HEADINGS = ["## Decisions", "## Action items"]   # the fixed format
REQUIRED_FACTS = ["3 November", "40,000"]                  # launch date, new budget

def test_planning_summary():
    runs = [summarise(PLANNING_TRANSCRIPT) for _ in range(3)]  # 3 API calls: costs credits
    for text in runs:
        assert all(h in text for h in REQUIRED_HEADINGS)       # the format holds every time
        assert all(f in text for f in REQUIRED_FACTS)          # the facts survive
        assert len(text.split()) <= 300                        # the length rule
    # assert runs[0] == SAVED_SUMMARY   <- the old test: wording varies, so it proves nothing

5.1.5 Thinking, effort and fast mode: how hard and how fast Claude works

The third complaint is the shallow plan. "What should we do first?" needs weighing: a budget cut, two deadlines and three teams competing for one designer. To keep the app snappy, Dalia had set every request to the lowest effort. Two request options decide how much work goes into a reply, and a third changes only its speed.

Thinking lets Claude reason before it answers, in thinking blocks that arrive ahead of the answer text. That is what a plan needs: room to weigh options before committing to the first sentence. There are two modes. Extended thinking is the older, manual one: thinking: {"type": "enabled", "budget_tokens": N} sets a thinking budget (at least 1,024 tokens and less than max_tokens), and Claude thinks before every answer. It is deprecated on Claude Opus 4.6 and Claude Sonnet 4.6, and Claude Opus 4.7 and later models reject it with a 400 error.

Adaptive thinking, set with thinking: {"type": "adaptive"}, lets Claude decide on each request whether to think and how much, so a simple question may get no thinking block at all. In Anthropic's internal evaluations it reliably performs better than extended thinking. It is on by default on Claude Sonnet 5.5, always on for Claude Opus 5.5, and off until you set it on Claude Opus 4.6 to 4.8 and Claude Sonnet 4.6.

Effort steers how many tokens Claude spends on the whole reply: text, tool calls and, with adaptive thinking, how readily and deeply it thinks. You set output_config: {"effort": ...} to low, medium, high, xhigh or max. The levels on offer vary by model (Claude Haiku 4.5 has no effort setting), and the default is high on most models and medium on Claude Opus 5.5. Effort is guidance, not a budget: at low, Claude skips thinking on simple requests and thinks less on hard ones, while max_tokens stays the hard ceiling.

Fast mode is different in kind: the same model, served with a faster configuration, writing up to 2.5 times more output tokens per second at premium prices, with no change in intelligence. You opt in with speed: "fast" and a beta header. It is a research preview, with access on request, on Anthropic's own API for Claude Opus 5.5, Opus 5 and Opus 4.8 only. It speeds up the stream once it starts, not the wait for the first token.

Picture briefing a consultant. Thinking is time at the whiteboard before they answer. Effort is how many hours you authorise for the job. Fast mode is a faster typist: the same answer, written up sooner.

Option What it changes Reach for it when
Adaptive thinking Whether and how much Claude reasons before answering, decided per request Hard multi-step questions, on models that support it
Extended thinking A manual thinking budget on every request (budget_tokens) Only on older models without adaptive thinking, such as Claude Haiku 4.5
Effort Tokens spent on the whole reply, thinking included Lower for routine or latency-sensitive calls, higher for hard reasoning
Fast mode Output speed of the same model, at a premium price A supported Opus model is the right choice but its output must arrive faster

Memorise the middle column: thinking and effort change depth, fast mode changes only speed. Dalia's fix follows. Look at the effort line, which picks depth per kind of question, and at the roomy max_tokens, which leaves space for thinking plus the answer.

import anthropic

client = anthropic.Anthropic()

def ask(question: str, transcript: str, hard: bool):
    return client.messages.create(
        model="claude-sonnet-5-5",            # adaptive thinking is on by default
        max_tokens=16000,                     # room for thinking AND the answer
        output_config={"effort": "high" if hard else "low"},   # depth per kind of question
        messages=[{"role": "user",
                   "content": f"<transcript>{transcript}</transcript>\n\n{question}"}],
    )

reply = ask("Given the budget cut and both deadlines, what should we do first?",
            transcript, hard=True)
print(reply.stop_reason, any(b.type == "thinking" for b in reply.content))  # did it think?

With planning questions at high, Claude thinks before answering and weighs the deadlines against the budget, while quick questions such as "who owns the pricing page?" stay at low and stay fast.

5.1.6 Zero-shot, single-shot and multi-shot prompting

Rohan has one more fair point: the layout of the summaries wanders. Action items come as bullets, then as a table, then with the owner last, and his task-tracker import breaks on every variant. Here the lever is the prompt itself, and in particular whether it shows Claude examples of the output.

  • Zero-shot: instructions only, no example. It is the right place to start, and for many tasks it is enough.
  • Single-shot (also called one-shot): one example of the output you want. It shows a format better than any description, but Claude learns from everything the example shows, so one example also teaches its accidents. If the only sample summary gives every action item a Friday deadline, many summaries will too.
  • Multi-shot (also called few-shot): several examples. Anthropic's prompting guide counts examples among the most reliable ways to steer format, tone and structure. It recommends 3 to 5 examples that are relevant (close to the real task) and diverse (varied enough, edge cases included, that Claude does not pick up unintended patterns). It also wants them structured: each in <example> tags, all inside <examples>, so Claude can tell them from instructions.

Here is Dalia's multi-shot prompt, each example abbreviated to a note of its contents. Look at how the examples differ in size and edge case, and at the transcript placed first, as the guide recommends for long inputs.

<transcript>{transcript}</transcript>

Summarise the meeting in the transcript above. Follow the format of the examples exactly: a "Decisions" list, then an "Action items" list where each line reads OWNER | TASK | DUE. If a task has no owner or no date, write "unassigned" or "no date"; never invent one.

<examples>
<example>(a stand-up: one decision, two action items, each with an owner and a date)</example>
<example>(a client call: no decisions, three action items, one with no owner)</example>
<example>(a planning meeting: four decisions, five action items, one with no date)</example>
</examples>

Examples cost tokens on every call, so keep them short and representative. They do not switch sampling off either: the wording still varies, but the shape stops varying, and the shape is what the importer and the property test depend on.

5.1.7 The exam traps

Most traps in this skill pull the wrong lever: a sampling knob for a depth problem, a speed option for a quality problem, a bigger window for an output cap.

  • ✗ Setting temperature to 0 to get identical output. ✓ Expect variation and design for it. A temperature of 0 was never fully deterministic, and Claude Opus 4.7 and later models reject non-default values.
  • ✗ Changing the temperature to get deeper reasoning. ✓ Raise effort, with thinking available. Temperature changes how varied the word choice is, not how much Claude reasons.
  • ✗ Turning on fast mode to fix a weak answer. ✓ Fast mode is the same model writing faster at a premium price. It changes speed, never quality.
  • ✗ Treating a reply cut off mid-sentence as a context-window problem. ✓ Read stop_reason. max_tokens means your output cap was too small for the reply plus any thinking; an input that truly exceeds the window fails with a 400 error.
  • ✗ Running every request at maximum effort. ✓ Match effort to the request: low for routine, latency-sensitive calls and higher for hard reasoning. Effort is guidance; max_tokens is the hard cap.
  • ✗ Fixing format drift with one example, or with several near-identical ones. ✓ Use 3 to 5 relevant, varied examples in <example> tags, so Claude learns the format rather than one example's accidents.

Four tempting knobs for a shallow answer, one right one

Raise temperaturevaries the words, not the reasoning
Lower temperaturerejected by new models, and no deeper
Turn on fast modesame model, faster output
Max effort on every callcost and latency everywhere
Raise effort for hard questionsClaude thinks more readily and more deeply
Temperature, fast mode and blanket maximum effort each change something other than the depth of this particular answer; raising effort where the question is hard does.

5.1.8 Put it together: tune a summariser and break it

You now have every piece: tokens and next-token generation, the window and the output cap, sampling, the depth and speed options, and examples. Build a small version of Dalia's summariser and watch each of her three levers move.

Everything else in Domain 5 builds on these units. Technical fundamentals (5.2) cover the SDKs and connections under every call. Model selection (5.3) decides which model handles which request, including which models offer adaptive thinking. Cost and token management (5.4) turns tokens into money through budgets, tracking and caching. And prompt engineering (6.2) takes examples further, alongside clarity, placement and iteration.

Key takeaways

  • ✓ Claude writes one token at a time, each picked from probabilities given everything before it; tokens measure size, cost and speed, and their count depends on the model's tokenizer.
  • ✓ The context window holds input and output together, and max_tokens caps the reply with thinking spending from it, so a max_tokens stop means your cap was too small, not the window.
  • ✓ Sampling makes the same prompt give different wording; no setting guarantees identical output, and Claude Opus 4.7 and later models reject non-default temperature, top_p and top_k.
  • ✓ Design for variation: fix the format, show examples and test the properties of a good output rather than its exact text.
  • ✓ Adaptive thinking lets Claude decide per request whether to reason first, extended thinking's manual budget_tokens belongs to older models, and effort is the main lever for depth, cost and latency.
  • ✓ Fast mode makes supported Opus models write up to 2.5 times faster at a premium price, without changing their intelligence.
  • ✓ Start zero-shot; when format or conventions drift, use 3 to 5 relevant, varied examples in <example> tags, because a single example gets copied in detail.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

27 CCDV-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 27 questions on Model Selection and Optimization alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources