Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 2 › Lesson 2.1

CCAR-P · Domain 2 · 13% of the exam · Lesson 2.1 · 22 min read

Selecting Claude models: trade-offs, constraints and the evidence to decide

How an architect picks a Claude model per workload: the tiers compared, platform, residency and lifecycle limits, and evals on your own data that decide.

Written against objective 2.1 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.1.1 Why one solution rarely runs on one model

Portolan Travel manages business travel for about 400 corporate clients, and three Claude workloads are heading for production. The traveller assistant answers people mid-trip ("my connection is cancelled, what now?") and rebooks them through Portolan's booking tools. Receipt extraction reads about 50,000 photographed and scanned expense receipts a day and turns each into fields: merchant, date, amount, currency, tax and category. Policy exceptions help an approver decide whether a request that breaks a client's travel policy, such as a business-class seat on a six-hour flight, can be approved, and why.

The pilot ran all three on Claude Opus 5.5, and the design review turns into a tug of war. The engineering lead wants Opus everywhere, because Anthropic's documentation names it the starting point for most workloads. Rasmus, the finance director, has multiplied the receipt volume by Opus prices and wants the cheapest model everywhere. The head of traveller services wants whichever model answers fastest. Each of them is right about one workload and wrong about the other two.

Folasade, Portolan's solution architect, changes the question. "Which model do we use?" has no answer at company level; it has one answer per workload, and each answer must pass three tests. It must fit the deployment's hard constraints: what the model can hold, where it may run, how long it will be offered. It must strike the balance of quality, speed and cost this workload needs. And it must hold up in an evaluation on Portolan's own data. That discipline is model selection.

2.1.2 What separates the current models

Drop one belief first: that the tiers are "good, better, best" and the only question is budget. They differ on several axes at once, and some are hard limits rather than quality. Anthropic's models overview lists four current models. Prices are per million tokens, the word pieces a model reads (input) and writes (output). The context window is how many tokens one request can hold, and maximum output caps what one reply can write.

Model What Anthropic positions it for Speed; price per million input / output tokens Limits and reasoning controls
Claude Fable 5.1 The most capable model open to all customers: demanding reasoning, long-horizon agentic work Slower; $10 / $50 1M context, 128K output; adaptive thinking always on; default effort high
Claude Opus 5.5 Long-running agentic coding and knowledge work; the docs' starting point for most workloads Moderate; $4 / $20 1M context, 128K output; adaptive thinking always on; default effort medium
Claude Sonnet 5.5 The best combination of speed and intelligence: everyday code, analysis, content, agentic tool use Fast; $2 / $10 1M context, 128K output; adaptive thinking; default effort high
Claude Haiku 4.5 The fastest model, with near-frontier intelligence: real-time, high-volume and sub-agent tasks Fastest; $1 / $5 200K context, 64K output; extended thinking with a manual budget; no effort parameter

Memorise the order of the tiers and the kinds of limit. Prices and version numbers change with each release, so check the models overview before you quote them.

The last column names the reasoning controls. With adaptive thinking, the model decides how much to reason before it answers, and the effort parameter steers how many tokens it spends on thinking, tool calls and the reply. Haiku 4.5 offers only the older extended thinking, where you set a thinking budget yourself. These controls come with the model; tuning them is a separate decision. Two things separate nothing: every current model reads images and uses tools.

Read the price column with care. Every model in the table except Haiku 4.5 uses a newer tokenizer that turns the same text into about 30% more tokens, so price per token understates the gap. The honest comparison is cost per completed task.

Folasade's first pass uses the table only to rule things out. Receipts are images, which every model reads, so vision decides nothing. For the largest clients, the policy-exceptions prompt carries a full policy book, client overrides and a year of trip history, up to about 400,000 tokens. Haiku 4.5's 200K window removes it from that route before any quality test.

2.1.3 One solution, several models: route by workload

Here is the question that trips people up: if the docs name Opus 5.5 as the place to start, why not stay there? Because a starting point for testing is not a destination for every workload. Anthropic's model guide describes two ways to begin. Capability-first starts on a strong model, proves the quality bar, then moves work to cheaper settings or tiers where the evaluation allows. Efficiency-first starts on a fast, low-cost model and upgrades only for a proven gap; it suits prototypes, tight latency targets and high-volume, straightforward tasks.

Which tier to test first depends on the one requirement each workload cannot give up, and Folasade writes it down for all three.

The requirement that decides each Portolan workload

Traveller assistant

A traveller at the gatemulti-turn, calls booking tools
Decided bylatency, with reliable tool use
Start on Sonnet 5.5

Receipt extraction

50,000 a dayshort, structured, checkable
Decided bycost per receipt at a fixed accuracy bar
Start on Haiku 4.5

Policy exceptions

About 300 a dayan approver can wait a minute
Decided byquality of judgement; errors cost money
Start on Opus 5.5
Each workload has one requirement it cannot trade away, and that requirement, not the model's rank, picks the first tier to test.

A starting tier is a hypothesis, not a decision. The evaluation measures each candidate against a bar that a capable model sets, then confirms the tier or moves it. The architect's other choice is how the models combine inside one solution, and there are four common shapes.

Pattern When it wins What it costs
One model for everything A single workload, or several of uniform difficulty; an early prototype Overpays on the easy work or underperforms on the hard work; no route can be tuned alone
A model per workload Workloads differ in the requirement that decides them, as at Portolan A route table, plus an eval set and monitoring per route
Escalation inside a workload Mostly routine traffic with a hard tail you can detect, such as a failed validation A reliable escalation signal, and two models to evaluate on one path
Two models in one agent An agent that is routine with a few hard decisions, or work that fans out into independent pieces More design and testing; the pair must beat one model swept across effort levels, which Anthropic's guide sets as the baseline

Portolan uses the second and third rows. Each workload gets its own route, and any receipt whose extracted total disagrees with the card transaction goes to Sonnet 5.5 before a person sees it. The check is cheap and exact, which is what makes escalation safe here.

2.1.4 Where the model can run: platforms and data residency

A model that wins on quality and cost is still the wrong choice if it cannot run where your contracts say data must be processed. Claude is offered on several platforms, and their differences remove options rather than rank them, so check them per model before any trade-off.

Platform Who runs it; whose retirement dates apply How you keep inference in one geography
Claude API Anthropic; Anthropic's dates inference_geo: "us" per request or as a workspace default, on Claude 4.6 and later models, at 1.1 times the standard price
Claude Platform on AWS Anthropic, billed through AWS Marketplace; Anthropic's dates inference_geo, as on the Claude API
Microsoft Foundry Anthropic, billed through Azure Marketplace; Anthropic's dates A US Data Zone Standard deployment, for models that offer one
Amazon Bedrock AWS; AWS sets its own dates A regional endpoint or geographic inference profile instead of the global endpoint
Google Cloud Agent Platform (Vertex AI) Google; Google sets its own dates A regional or multi-region (us, eu) endpoint instead of the global one

Memorise the pattern, not the rows: on Anthropic-operated platforms a parameter or deployment type pins the geography, and on partner platforms the endpoint you call decides it. Either way, pinning a geography costs about 10% more on current models.

On the Claude API, the default "global" lets inference run in any available geography. Today the only other value is "us", and each response's usage reports where inference actually ran. Setting allowed_inference_geos: ["us"] and default_inference_geo: "us" on a workspace is the preventive version of the control: a request that omits the parameter runs in the US, and one that asks for another geography is rejected. Storage is a separate setting: the workspace geo, fixed when the workspace is created, decides where its data rests.

Features differ too. Anthropic's pages for Bedrock, Google Cloud and Foundry all list the Message Batches API as unsupported, for example, so check every feature your design depends on against the chosen platform.

Portolan's contracts allow global processing today, but a prospective client, an aerospace manufacturer, asks in its security questionnaire for US-only processing. Folasade checks each route. The assistant and the exceptions route can add inference_geo: "us" and pay 10% more. The receipt route cannot: Haiku 4.5 rejects the parameter with a 400 error. For that client's receipts she can keep Haiku 4.5 on a US-only deployment (a US endpoint on Bedrock or Google Cloud, or a US Data Zone deployment on Foundry) or move to Sonnet 5.5 on the Claude API. The eval set will price both, and she records the consequence before sales promises anything.

2.1.5 Pinned IDs and the life of a model

What exactly does the model string in your configuration point at, and for how long? Two facts answer it.

First, every Claude model ID names a pinned snapshot. Anthropic does not change the weights behind an existing ID; an improved model ships under a new ID. It is the difference between ordering by part number and ordering "the current drill": the part number never changes under you. From the 4.6 generation on, IDs are dateless (claude-sonnet-5-5), and the dateless ID is the snapshot itself, not a pointer to the newest version. Older models have dated IDs (claude-haiku-4-5-20251001) plus a shorter alias on the Claude API (claude-haiku-4-5) that resolves to the most recent dated snapshot of that minor version. The pinning guarantee covers the ID, not the alias.

Second, every ID has a lifecycle, and the deprecations page tracks each one.

The lifecycle of a model ID

ACTIVEfully supported and recommended
LEGACYno longer updated; may be deprecated later
DEPRECATEDstill answers; replacement named, retirement date set
RETIREDevery request fails
A model moves from active to retired on a published schedule, and only the last step breaks requests, so the migration window is yours to use.

Anthropic notifies customers with active deployments at least 60 days before it retires a publicly released model, and warns that deprecated models are likely to be less reliable than active ones. The Console's usage export breaks usage down by API key and model, which is how you find every caller of a deprecated ID.

The lifecycle is part of the choice. The deprecations page gives every active model a "not sooner than" retirement date, and Haiku 4.5, a generation older than the rest of the lineup, has by far the earliest. That does not disqualify it, because notice comes at least 60 days ahead. It does mean the receipt route will need a successor first, so its eval set must be ready to rerun. Look at the comments on Portolan's route table: the ID form each route uses, and the rule at the bottom.

# Hypothetical route table: one pinned model ID per workload
traveller_assistant:
  model: claude-sonnet-5-5            # 4.6+ generation: the dateless ID is the snapshot
  effort: low                         # latency first; tuned as its own decision
receipt_extraction:
  model: claude-haiku-4-5-20251001    # dated ID, not the claude-haiku-4-5 alias
  escalate_to: claude-sonnet-5-5      # totals that fail the card-feed check
policy_exceptions:
  model: claude-opus-5-5
  effort: high
change_rule: >-
  No model ID changes without a passing run of that route's eval set
  and a decision record. Check the deprecations page monthly.

2.1.6 Let your own evaluation decide, and decide again

A public leaderboard says one model scores higher, so why not use that? Because it does not run your prompts, your data or your edge cases. Anthropic's cost guide says the ranking between models flips by workload in ways no price list predicts. Its model guide calls a good evaluation set, run with your actual prompts and data, the most important step in deciding whether to change models.

An evaluation set (eval set) is a collection of real inputs with known good outcomes that you run through a model and score automatically where you can. For model selection, it shows what "good" looks like before you economise.

Setting the bar, then testing a cheaper model against it

SET THE BARcapable model on your eval set
CHALLENGEcheaper candidate, same set
COMPAREquality, hardest tenth, latency, cost per task
SHADOWwinner beside the current model on a slice of live traffic
SWITCHpin the ID, keep the suite running
SWITCH → SET THE BAR · rerun when a model ships or is deprecated
The capable model defines the quality bar on your own data; a cheaper model earns the route only by meeting it, and the loop restarts whenever the lineup changes.

Two details make the comparison honest. Compare cost per completed task, because a failed task still bills its tokens, then the retry, then the damage downstream. And price the tail: Anthropic's guide says to compare models on the hardest tenth of your tasks, where cheaper models fail and the bill is decided.

Portolan's receipt eval holds 600 real receipts with fields verified by the expense team, weighted like live traffic, including crumpled, handwritten and foreign-currency ones. Opus 5.5 sets the bar at 99.2% of fields correct. Haiku 4.5 scores 99.0% overall and 95% on the hardest tenth, against 97% for Opus. Most receipts it gets wrong fail the card-feed check and escalate, so the route as a whole meets the bar at a fraction of the cost per receipt, and Haiku wins it.

The other routes settle the same way. On 150 past exceptions graded by senior approvers, Opus 5.5 at high effort meets the bar and Fable 5.1 scores no better, so Fable stays a challenger. For the assistant, Sonnet 5.5 at low effort meets the latency target, while Haiku 4.5, faster still, drops more multi-step rebookings.

Folasade's evidence goes into a decision record, so the next architect can rerun the decision instead of guessing at it. Look at the last two lines: the known consequences and the triggers to revisit.

Decision: traveller assistant on claude-sonnet-5-5 (effort low); receipt extraction on claude-haiku-4-5-20251001, with receipts that fail the card-feed check re-run on claude-sonnet-5-5; policy exceptions on claude-opus-5-5 (effort high).
Evidence: eval sets of 400 conversations, 600 receipts and 150 graded exceptions, each weighted like live traffic with a hardest-tenth slice; two weeks in shadow on 5% of traffic.
Rejected: Opus everywhere (receipt cost with no measured quality gain); Haiku everywhere (missed multi-step rebookings; 200K window too small for the largest policy books); Fable 5.1 for exceptions (no gain on the graded set).
Known consequences: a US-only clause would move the receipt route to another platform or to Sonnet 5.5, because Haiku 4.5 does not accept inference_geo; Haiku 4.5 has the earliest retirement date in the lineup, so its successor is evaluated first.
Revisit when: a new model ships in any tier, a model in use is deprecated, or a route's weekly eval score falls below its bar.

Each release is a cue to rerun the suite, because nothing changes until you switch IDs. In Anthropic's measurements, each newer model solved at least as many tasks as the one before, usually for less per solved task. That is a reason to test promptly, not to switch blind: read the migration guide, fix what breaks, rerun the evals, shadow, and move route by route.

2.1.7 The exam traps

Every trap here picks a model by something other than the workload's requirement, the deployment's constraints and the evidence.

  • ✗ Sending every workload to the most capable model because it feels safest. ✓ Route per workload. On a high-volume, checkable task the top tier buys cost and latency, not quality you can measure.
  • ✗ Moving everything to the cheapest model to cut the bill. ✓ Move a workload down only when its eval, hardest tenth included, still meets the bar, and keep hard judgement on a capable model.
  • ✗ Choosing by public benchmarks or by price per token. ✓ Choose by your own eval set and cost per completed task. Rankings flip by workload, and tokenizers differ between models.
  • ✗ Treating a model ID or alias as "always the latest". ✓ IDs are pinned snapshots and new models ship under new IDs; aliases exist only for pre-4.6 models and stay within one minor version. Upgrades are deliberate, tested changes.
  • ✗ Checking residency and platform after the model is chosen. ✓ Filter first. inference_geo: "us" works only on Claude 4.6 and later models, and partner platforms have their own regions, features and retirement dates.
  • ✗ Waiting for the retirement notice, or switching on launch day. ✓ Rerun the evals when a model ships or is deprecated, and migrate well before the retirement date.

2.1.8 Put it together: pick and defend a model for one workload

You now have every piece of the method. The quickest way to make it stick is to run one workload through the method, hit a constraint on purpose, and write the record.

Later objectives build on this decision. Evaluating accuracy-latency trade-offs (3.3) tunes effort and related settings on the model you picked. Designing evaluation datasets (4.2) builds the eval sets this choice rests on. Diagnosing system issues (4.4) includes spotting in production that a workload has outgrown its model. Optimising token usage, latency and cost (4.5) turns cost per task into caching, batching and budgets.

Key takeaways

  • ✓ Choose a model per workload by the requirement that decides it: latency, cost at volume or quality of judgement.
  • ✓ The lineup runs from Claude Haiku 4.5 (fastest, cheapest, 200K context, no effort parameter) to Claude Fable 5.1 (most capable, slowest); hard limits rule models out, and trade-offs choose among the rest.
  • ✓ Route workloads to different models inside one solution, and escalate within a workload only when a cheap, reliable check spots the hard cases.
  • ✓ Residency and platform are filters: on the Claude API inference_geo: "us" works only on Claude 4.6 and later models, and on partner platforms the endpoint decides the region.
  • ✓ Model IDs are pinned snapshots, new models ship under new IDs, and every ID moves from active to retired with at least 60 days' notice.
  • ✓ Set the bar with a capable model, test cheaper candidates against it on your own eval set, hardest tenth included, and compare cost per completed task.
  • ✓ Record each choice with its evidence and triggers, and rerun the evals whenever a model ships or is deprecated.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

24 CCAR-P questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 24 questions on Claude Models, Prompting & Context Engineering alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources