Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 5 › Lesson 5.3

CCDV-F · Domain 5 · 16.8% of the exam · Lesson 5.3 · 21 min read

Choosing a model: capability, latency, cost and safe upgrades

How to match Haiku, Sonnet, Opus or Fable to each workload, which models think adaptively, and how to test a new model release before you switch.

Written against skill 5.3 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.3.1 Why one model for every job is the wrong default

Mireille leads engineering at Brieflark, a legal-tech startup that sells an assistant to small law firms. It runs three jobs on Claude. First, every incoming email gets a triage label (new client, court deadline, billing or noise), about 40,000 emails a day, sorted before a lawyer opens the inbox. Second, a lawyer asks for a contract clause drafted in the firm's house style and waits at the screen for it. Third, before a client signs an acquisition, the firm asks for a risk analysis across dozens of contracts: which ones give the other side rights when ownership changes, and which ones contradict each other. That job is rare and runs for minutes, and a missed clause can cost the client dearly.

The first version sent all three jobs to the most capable model, because "the best model" felt safe. Triage became the largest line on the bill, and its labels arrived after the lawyers had opened the emails. Teodor, who runs product, proposed the opposite: the cheapest, fastest model for everything. Triage stayed accurate, but the risk reports started missing clauses the bigger model had caught. Both plans made the same mistake. They chose one model for the company instead of one for each job.

There is no single best Claude model. The family comes in tiers that trade capability against speed and price, and the right tier depends on what a particular workload needs. Model selection is the habit of stating that need (how good, how fast, how cheap), starting from the tier that meets it, and proving the choice with an eval. An eval (short for evaluation) is a set of real inputs with known good answers that you run through a model and score. Model selection is also never finished. A new model has just been released, and Brieflark has to decide whether to move to it.

5.3.2 What each tier is built for

Here is the belief that trips people up: surely the most capable model is always the better choice? On hard problems it usually does give better answers. But capability is only one of three things you buy with every request, and the tiers exist because speed and price move the other way. A law firm works the same way. Nobody sends the senior partner to sort the morning post, and nobody asks the post room to review an acquisition.

Anthropic's documentation currently lists four tiers. The exam guide names three of them: Haiku, Sonnet and Opus. The fourth, Fable, is the newest tier and sits above Opus. Prices are per million tokens, the word pieces a model reads (input) and writes (output). Agentic work means Claude calling tools in a loop, step after step, until a multi-step task is done.

Tier (current model) What the docs position it for Speed; price per million input / output tokens
Haiku (Claude Haiku 4.5) Real-time apps, high-volume processing, sub-agent tasks; the fastest model, with near-frontier intelligence Fastest; $1 / $5
Sonnet (Claude Sonnet 5.5) Everyday code generation, data analysis, content creation and agentic tool use; the best mix of speed and intelligence Fast; $2 / $10
Opus (Claude Opus 5.5) Long-running agentic coding and knowledge work, large refactors, complex systems; the docs' starting point for most workloads Moderate; $4 / $20
Fable (Claude Fable 5.1) The highest capability: agent sessions that run for hours, deep multistep research, demanding reasoning Slower; $10 / $50

Memorise the order and the roles, not the numbers, which change with each release. Capability also includes hard limits. Claude Haiku 4.5 has a 200K-token context window (the most text it can consider at once) and writes at most 64K tokens, while the other three have 1M and 128K. All four read text and images and can use tools.

Now map Brieflark's jobs. Triage is short, well defined, huge in volume and easy to check against a known label: a Haiku job. Clause drafting needs quality while a person waits: a Sonnet job. Risk analysis reads dozens of contracts at once, which can outgrow Haiku's smaller window, and its misses are costly: an Opus job, with Fable in reserve if the evals show Opus falling short.

5.3.3 Quality, latency and cost: name the requirement first

"Which model is best for us?" has no answer until you say which of three things a workload cannot give up. Quality is how often the output must be right, and what a miss costs. Latency is how long a person or a downstream system waits for the answer. Cost is the price per request multiplied by the volume. The three pull against each other: a more capable tier is usually better on hard tasks, and slower and dearer per token.

So for each workload, find the requirement that decides. Brieflark's three jobs land on three different ones.

The requirement that decides each workload

Email triage

40,000 a daya label within seconds
Decided bycost and latency
Start on Haiku

Clause drafting

A lawyer is waitingand signs the result
Decided byquality, at interactive speed
Start on Sonnet

Risk analysis

A few a weekminutes are fine
Decided byquality; misses are costly
Start on Opus
Each workload has one requirement it cannot give up, and that requirement, not the model's rank, picks the starting tier.

Two refinements keep this honest. First, compare models on cost per completed task, not on the price per token. A more capable model often finishes with fewer turns, less re-reading and fewer retries, while a failed task still bills its tokens, and then the retry. Anthropic's cost guide adds a sharper rule: compare on the hardest tenth of your tasks, because that is where a cheaper model fails and where the bill is decided. Second, the tier is not the only dial. The effort parameter trades intelligence for speed and cost inside one model, and the docs note that tuning effort is often a better lever than switching models.

5.3.4 Efficiency-first or capability-first

Knowing the requirement tells you where you are heading, not where to start testing. Start at the top and you rarely find out that a smaller tier would do; start at the bottom and you may blame your prompts for a capability gap. Anthropic's model guide describes two starting strategies.

Efficiency-first starts with a fast, low-cost model such as Haiku. You build, test thoroughly against your requirements, and upgrade only where the tests show a specific capability gap. It suits prototypes, tight latency targets, cost-sensitive products and high-volume, straightforward tasks.

Capability-first starts with the strongest sensible model, which the docs currently name as Opus. You tune your prompts for it and check that it meets the bar, then look for savings: lower the effort, or move parts of the work down a tier where the evals hold. If the evals at the highest effort levels still fall short on demanding reasoning or long agent runs, you move up to Fable. It suits complex reasoning, nuanced judgment, advanced coding and work where accuracy outweighs cost.

Two ways to start

Efficiency-first

Start on Haiku
Test against the requirement
Upgrade only for a proven gap

how Brieflark chose for triage

Capability-first

Start on Opus
Prove the quality bar
Lower effort or tier where evals hold

how Brieflark chose for risk analysis

Both strategies end at the same place, an eval on your own data; they differ only in the direction you move from the starting tier.

Brieflark uses both. Triage went efficiency-first: Haiku hit the accuracy target on 500 past emails with lawyer-confirmed labels, so it never moved. Risk analysis went capability-first. Opus found every planted ownership-change clause in a test pack of contracts, and only then did Mireille try a lower effort level to see how much quality it cost. Drafting started on Sonnet, the docs' pick for everyday content creation, and stayed there.

5.3.5 Adaptive thinking: which models think for themselves

The risk analysis needs Claude to reason before it answers, so can you just switch reasoning on? That depends on the model. With thinking, Claude works through a problem in thinking blocks before its answer, which helps on analysis, coding and long agent tasks. Thinking tokens are billed as output tokens and count toward max_tokens, the cap you set on each reply, even when the thinking text is hidden from you.

Adaptive thinking lets the model decide when and how deeply to think on each request, steered by the effort level. The older mode, extended thinking, has you set a thinking budget yourself: a token target in budget_tokens on every request. Think of telling a colleague how much care a job deserves rather than handing them a stopwatch; they can then spend a minute on an easy question and an hour on a hard one. Which of the two a model supports is part of choosing it, because the tiers differ.

Model Thinking What your request controls
Claude Fable 5.1, Claude Opus 5.5 Adaptive, always on effort (default high on Fable 5.1, medium on Opus 5.5); a thinking budget or "disabled" returns a 400 error
Claude Sonnet 5.5 Adaptive, on by default effort (default high); a budget or "disabled" returns a 400 error, and "between_tools" turns off the thinking before the answer
Claude Haiku 4.5 Extended thinking only, off by default A manual budget_tokens; "adaptive" returns a 400 error and there is no effort parameter

Remember the pattern rather than every cell: the newer tiers think adaptively and you steer them with effort, while Haiku 4.5 does not support adaptive thinking at all.

For Brieflark this sorts itself out. Risk analysis on Opus always thinks, which is the point, and Mireille tunes its depth with effort. Triage on Haiku runs without thinking, because a four-way label does not need it and every thinking token would be billed 40,000 times a day. Look at ROUTES, which pins one model and effort level per workload, and at the last line, which picks the reply by block type because a thinking model's reply can start with a thinking block.

import anthropic

client = anthropic.Anthropic()

ROUTES = {  # one pinned model per workload, each chosen by evals
    "triage": {"model": "claude-haiku-4-5-20251001"},  # no effort parameter on Haiku 4.5
    "drafting": {"model": "claude-sonnet-5", "effort": "medium"},  # the previous Sonnet release
    "risk": {"model": "claude-opus-5-5", "effort": "high"},  # set explicitly: the default is medium
}

def run(workload: str, prompt: str) -> str:
    route = ROUTES[workload]
    extra = {"output_config": {"effort": route["effort"]}} if "effort" in route else {}
    response = client.messages.create(
        model=route["model"], max_tokens=16000,
        messages=[{"role": "user", "content": prompt}], **extra,
    )
    return next(b.text for b in response.content if b.type == "text")  # by type, not content[0]

5.3.6 A new model release is a change you test

"The new model scores higher on every benchmark, so switching is a one-line change." The line is one line; the change is not. Start with the fact that makes testing possible: every Claude model ID names a pinned snapshot, and a new model ships under a new ID. The model behind your ID does not change, so the upgrade is yours to make and yours to test. It works like a pinned library version in a lockfile: nothing moves until you bump it, and before a major bump you read the changelog and run the tests. Here the changelog is the new model's migration guide, and the tests are your evals.

You cannot stay pinned forever, though. Anthropic retires older models: a deprecated model still answers but has a named replacement and a retirement date, and once it is retired, every request to it fails. The model deprecations page lists each model's status, so every workload moves eventually, as a staged release like any other change.

Breaking behaviour changes between releases come in three kinds, and only the first one announces itself.

Kind of change Recent examples How you catch it
Requests that now fail Claude Sonnet 5.5 returns a 400 error for five settings that older models accepted, among them a thinking budget, a non-default temperature, a forced tool_choice and thinking: {"type": "disabled"} The migration guide, then one test run that fails fast
Responses with a new shape A request with no thinking field now thinks on Sonnet 5.5 (it did not on Sonnet 4.6 or Haiku 4.5), so a reply can open with a thinking block and code that reads content[0].text breaks; text between tool calls can arrive inside thinking blocks Parsing tests, and reading blocks by type
Silent shifts in behaviour and cost Sonnet 5.5's effort levels are recalibrated; Opus 5.5 defaults to medium effort where Opus 5 ran at high; the same text makes about 30% more tokens on Sonnet 5.5 than on Sonnet 4.6 or Haiku 4.5 Evals on quality, latency and cost per task, old model against new

The third row is the dangerous one. Nothing errors, the dashboards stay green, and the output has quietly changed: a prompt tuned for one model can land differently on the next, and a default you never set has moved.

Deciding on an upgrade uses the same method as choosing a tier, with the migration guide read first. The docs call a good evaluation set, run with your actual prompts and data, the most important step.

Testing a new release before you switch

READthe migration guide
FIXsettings that now fail
RUNyour evals on old and new
COMPAREquality, latency, cost per task
SWITCHper workload, or stay pinned
The new model earns production traffic workload by workload, and only after it beats the pinned model on your own evals.

Brieflark's drafting route runs on the previous Sonnet, Claude Sonnet 5. The real service also sends two settings the sketch above leaves out: thinking: {"type": "disabled"} for quicker replies, and a forced tool_choice so every draft arrives through an insert_clause tool. Claude Sonnet 5.5 rejects both. The guide's fix is between_tools, the new model's lowest thinking setting, and tool_choice set to auto. The tool is marked strict: true so its input still matches the schema, and the prompt says when to call it. Then comes the quiet part: effort is recalibrated, so Mireille reruns the drafting evals at two effort levels on both models before moving the route. Triage and risk analysis get their own evals.

5.3.7 The exam traps

Every trap here picks a model by something other than the workload's requirement and the evidence. The fix is almost always to route per workload and let evals decide.

  • ✗ Sending every workload to the most capable model because it feels safest. ✓ Match the tier to each workload. On 40,000 short labels a day, the top tier buys cost and latency, not quality you can measure.
  • ✗ Moving everything to the smallest model to cut the bill, whatever happens to quality. ✓ Move a workload down only where evals show it still meets its quality bar. Deep reasoning and multi-step tool use are where smaller tiers slip first.
  • ✗ Comparing models by the price per token. ✓ Compare cost per completed task on your own traffic, hardest tasks included. Failures and retries bill too, and a stronger model often does less work per task.
  • ✗ Changing tier when a setting would do. ✓ Sweep effort on the current model first; the docs call it often a better lever than switching models.
  • ✗ Assuming thinking works the same on every model. ✓ Check support per model. Fable 5.1, Opus 5.5 and Sonnet 5.5 think adaptively and reject a thinking budget or "disabled"; Haiku 4.5 has only manual extended thinking.
  • ✗ Treating a new release as a drop-in upgrade because its benchmarks are better. ✓ Read the migration guide, fix what now fails, and run your own evals on old and new before moving each workload.

5.3.8 Put it together: route a workload and test an upgrade

You now have the whole decision. You know what each tier is for, which requirement settles a workload, where to start testing, which models think adaptively, and why a release is a change you test. The quickest way to make it stick is to run one workload on two models, watch a release break it, and let the numbers decide.

The rest of the guide builds on this decision. Cost and token management (5.4) turns the cost column of your evals into a budget, with usage tracking, cost models and prompt caching. Prompt engineering (6.2) is what you revisit on every model change, because a prompt tuned for one model can land differently on the next. Output handling (6.3) covers defensive parsing, such as reading content blocks by type, so a new response shape cannot break your code.

Key takeaways

  • ✓ There is no single best model: choose one per workload by what that job needs in quality, latency and cost.
  • ✓ Haiku is the fastest and cheapest tier, Sonnet balances speed and intelligence, Opus handles complex long-running work, and Fable, the newest tier, sits above Opus.
  • ✓ Compare candidates on cost per completed task on your own traffic, and try effort on the current model before you change tier.
  • ✓ Efficiency-first starts small and upgrades for proven gaps; capability-first starts strong and optimizes down; both end in an eval on your own data.
  • ✓ Claude Fable 5.1 and Claude Opus 5.5 always think adaptively, Claude Sonnet 5.5 does by default, and Claude Haiku 4.5 offers only manual extended thinking.
  • ✓ Model IDs are pinned snapshots and retired models stop answering, so every upgrade is a change you make: read the migration guide, fix what fails, and run your evals before you switch.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

27 CCDV-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 27 questions on Model Selection and Optimization alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources