Home › Study guides › CCAO-F › Domain 3 › Lesson 3.3
CCAO-F · Domain 3 · 12% of the exam · Lesson 3.3 · 20 min read
Matching the model to the task: cost, speed and quality
How to weigh cost, speed and quality when choosing between Haiku, Sonnet and Opus, test the choice on real cases, and mix models across one workflow.
Written against objective 3.3 of the official CCAO-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.3.1 Why one model for every job costs you twice
It is the second week of January at Aldermoor University, and the admissions office is at full stretch. Harriet runs admissions operations with a team of six advisers, and three kinds of work land on them every day. About 300 enquiry emails ask the same few things: when is the deadline, which documents are needed, has my transcript arrived. About 40 scholarship decisions each need a letter that states the award exactly and reads as if it were written for that one student. And every week brings four or five appeals, each with a stack of documents and five working days for Harriet to send the appeals committee a careful recommendation.
At the start of the season the team agreed to run everything on the most capable model, because quality matters. The enquiry drafts were excellent, and slow to arrive. By mid-afternoon several advisers had hit their usage limits with a hundred emails still waiting. So the team switched to the fastest model for everything, and the next day the queue cleared by lunch. Then an appeal summary arrived. It listed all five documents neatly but never connected the dates on a doctor's letter to the exam where the student's grade dropped, the one point the appeal turned on.
Neither model failed. The mistake was one model for three very different jobs. Claude comes in tiers, from the fast, low-cost Haiku through the balanced Sonnet to Opus, the most capable of the three, and each strikes a different balance of cost, speed and quality. Model selection is the skill of starting from what a task needs and choosing the model that meets that need without spending more time or usage than it must.
One model for three jobs
Most capable for everything
Fastest for everything
Matched to each job
3.3.2 Cost, speed and quality pull against each other
It's tempting to treat the most capable model as the safe default for everything, and exam questions are built around that temptation. It fails because the three things you want from a model trade against one another. Deeper reasoning takes longer and uses more of your allowance. Speed and low cost come with less depth on a hard judgment. No model gives you the most of all three at once.
Tradespeople have an old saying for this triangle: fast, cheap, good, pick two. A kitchen fitted in three days at a bargain price is rarely fitted well. Model selection is the same bargain. You decide which of cost, speed and quality a particular task can least afford to lose, and you accept a trade on the other two.
That decision starts with the task, not the model menu. Four questions do most of the work: how much of it there is (VOLUME), how soon it is needed (DEADLINE), how much reasoning and judgment it takes (COMPLEXITY), and what a mistake would cost (STAKES). Here are Harriet's three jobs run through them.
| Question | Enquiry replies | Scholarship letters | Appeal recommendations |
|---|---|---|---|
| Volume | About 300 a day | About 40 a day | Four or five a week |
| Deadline | Same day | Within the week | Five working days |
| Complexity | Low: answers come from the published guide | Medium: exact figures in a personal tone | High: several documents weighed against policy |
| Stakes | Low per email, and easy to check | Medium: a wrong amount is a real problem | High: a student's place depends on it |
| Optimise for | Speed and cost | A balance of all three | Quality |
Memorise the four questions; the cells are only Harriet's answers. Notice how the questions combine. A single enquiry is low-stakes and easy to check against the published guide, but 300 a day make speed and cost decisive. Appeals are so few that their cost barely registers, while a student's place depends on the quality.
3.3.3 In the apps, cost shows up as usage
In the Claude apps you don't see a price on each reply, so cost can sound like somebody else's concern. It isn't. On most plans it takes the form of usage limits: how much you can use Claude in a rolling five-hour session, with weekly limits on top on paid plans. How far that allowance stretches depends on the length and complexity of your conversations, the features you use, the model you choose and its effort level. Effort is a setting for how much thinking Claude puts into each reply. When you reach a limit, you wait for it to reset, move to a higher plan or, on paid plans, turn on usage credits and pay for the extra.
What draws on your usage
Model choice is a big part of that. Anthropic's Academy guide to choosing a model ranks the tiers by how hard they draw on your limit: Haiku lightest, Sonnet moderate, Opus heavy. Using Opus on work a lighter model could handle, it warns, costs usage for no gain and slows you down. That is how Harriet's advisers ran dry: 300 routine enquiries a day went through a model built for hard judgment. On a usage-based Enterprise plan the effect moves to the bill: there is no fixed allowance, and the organisation pays for what it uses.
Cost also runs the other way. A model that is cheap per reply is not cheap if every draft needs ten minutes of rewriting, because staff time is part of the cost too. So count the cost of a finished, sendable piece of work, not of one reply. And when Claude Developers build an automated workflow that connects to Claude through the API (application programming interface), cost becomes a price per token. A token is a small chunk of text, and each tier up costs more per token.
3.3.4 Start small, or start strong
Knowing the trade-offs still leaves a practical question: which model do you try first? Anthropic's guide to choosing a model describes two starting strategies, and each suits a different kind of task.
Efficiency-first means starting with the fastest, lowest-cost model, testing it on the real work, and moving up a tier only where it falls short. It fits high-volume, straightforward tasks where speed and cost matter, which describes the enquiry replies. Capability-first means starting with the most capable model, getting the quality right, and only then trying a smaller model or a lower effort setting once the task is proven. It fits complex reasoning and nuanced work where accuracy outweighs cost, which describes the appeals.
Two ways to find the right model
Efficiency-first
Capability-first
The scholarship letters show efficiency-first working as intended. Harriet started them on Haiku, with a brief that already held each student's course, award and achievements. The facts came out right, but the letters read like a form, with the same lines of congratulation for every student. She moved them up one tier to Sonnet, and the letters passed. She didn't jump to Opus, because Sonnet already met the bar and every step up costs time and usage on 40 letters a day.
The appeals went the other way. A missed connection can cost a student a place, so Harriet started them on Opus and checked its recommendations against five past appeals whose outcomes she knew. Once those were reliably right, she tried Sonnet on the same five. It summarised each document well but, in two cases, missed a link between documents, so the appeals stayed on Opus.
3.3.5 Test on real cases before you decide
Both strategies rest on one word: test. It's tempting to choose by reputation, or by one impressive answer to a made-up question. Neither tells you how a model handles YOUR work, with your formats, your awkward cases and your standards. Anthropic's model guide says to test with your actual prompts and data, comparing accuracy, quality and the handling of unusual cases before weighing cost. Claude 101, Anthropic's introductory course, describes a simple routine for your own work. Collect 5 to 10 past examples of a task you do regularly, prompt Claude to produce the same thing, and compare its output with what you know was right. Run that routine on two models and you have a model test.
Here is the test behind the scholarship letters' move from Haiku to Sonnet. Look at the mix of cases, which includes the awkward ones, and at the decision rule on the last line, written before any result came in.
Model test: scholarship decision letters
Cases: eight past decisions, with the letter we actually sent as the answer key. Three full awards, three partial awards with conditions, two declines.
Method: the same brief and the same decision record for every case, each model in a fresh chat.
For each draft, check:
- Every fact matches the decision record: name, award, amount, conditions, date to accept by.
- The tone fits the decision: warm for an award, respectful and clear for a decline.
- Nothing is added that the record does not say.
- Editing needed before it could be sent: none, light or heavy.
Decision rule: use the lowest-cost model whose drafts pass all eight cases with light edits at most.
A test tells you more than which model won. Look at where the failures fall. If every model makes the same kind of mistake, such as inventing a date the brief never gave, the problem is almost certainly the brief. Something is missing or unclear, and no model can supply a fact that isn't in front of it. Fix the brief and run the test again before you consider a bigger model. If the smaller model fails where the larger one passes on the same brief, you have found a real capability gap, and that is the reason to move up.
3.3.6 Mix models across one workflow
Harriet's office doesn't need one answer to "which model?" It needs one answer per step. Mixing models gives you the speed and low cost of a small model on the bulk of the work and the depth of a large one where it counts. Anthropic's model guide describes the same pattern for teams building on Claude: a lower-cost model does most of the work, and a more capable one takes the hard decisions. In the apps you can also change the model partway through a conversation, when one message turns out harder than the rest.
| Step | Model | Why this one | Check before it goes out |
|---|---|---|---|
| Routine enquiry replies | Haiku | Many, simple, checkable against the published guide | An adviser reads each draft; Harriet spot-checks a sample daily |
| Scholarship letters | Sonnet | Exact figures in a personal tone; passed the test | Every figure against the decision record |
| An "enquiry" that is really an appeal or a complaint | Step up to Sonnet or Opus | The task has changed from lookup to judgment | An adviser takes it over |
| Appeal recommendations | Opus | Several documents weighed against the appeals policy | Harriet checks every date; the committee decides |
The third row is the escalation rule. Move up a tier when the TASK changes, from routine processing to interpretation and judgment, not when the volume grows or someone feels uneasy. A busier day of deadline questions still needs only the fast model. One email that says "I want to challenge my decision" needs more, however quiet the day.
The last column matters as much as the second. A more capable model makes a good draft more likely; it doesn't turn the draft into a decision. The appeal recommendation still goes to the committee as a recommendation, and Harriet checks every date in it against the documents first. Model choice changes the odds, never who is accountable.
3.3.7 The exam traps
Questions on this objective describe a task and a model choice that doesn't fit it. The mismatch runs both ways, and a third kind of trap blames the model for a problem it didn't cause.
- ✗ Using the most capable model for everything, to be safe. ✓ Match the model to the task. The top model on routine volume costs time and usage and adds nothing the task needs.
- ✗ Using the fastest, lowest-cost model for nuanced, high-stakes work to save usage. ✓ Put a more capable model where reasoning and judgment matter. The saving disappears in rework and risk.
- ✗ Switching models to fix a prompt problem. ✓ Check the brief first. If every model makes the same kind of mistake, the missing piece is in the prompt.
- ✗ Choosing a model by reputation or one impressive answer. ✓ Test on a handful of real past cases with known good answers, then decide.
- ✗ Running a whole workflow on one model. ✓ Classify the steps, give each the model it needs, and step up only when the task changes.
- ✗ Treating the most capable model's output as final. ✓ Keep human review on high-stakes outputs, whichever model wrote them.
3.3.8 Put it together: choose a model for each step of real work
You now have every piece of the method, and the order matters as much as the pieces. A model picked before the task is understood is a habit; a model rolled out before it is tested is a guess. The quickest way to make the method stick is to test two models on your own work, and to see what a test is worth without its answer key.
From task to model, in order
The rest of Domain 3 turns to what happens inside a conversation. Managing context limits and memory (3.4) covers when a long chat should be restarted or summarised, which also changes how fast usage runs down. In Domain 7, optimising a whole workflow (7.3) builds on these per-step model choices. And if the office later wants every incoming enquiry drafted automatically, that is an integration to hand to Claude Architects and Developers.
Key takeaways
- ✓ Cost, speed and quality trade against each other, so start from the task: its volume, deadline, complexity and stakes.
- ✓ In the apps, cost is usage: the model, effort level, conversation length and heavy features all draw on your plan's limits, and more capable models draw more.
- ✓ Efficiency-first starts small and upgrades only where quality falls short; capability-first starts strong and steps down once the task is proven.
- ✓ Test candidate models on 5 to 10 real past cases with known good answers and a decision rule set in advance; a mistake every model makes points to the brief.
- ✓ Mix models across a workflow: a fast, low-cost model for routine volume and a more capable one for judgment, stepping up only when the task changes.
- ✓ No model choice removes the need for human review of high-stakes outputs.
Check your understanding
4 questions written for this lesson, then one from the CCAO-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
42 CCAO-F questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 42 questions on Product and Model Selection alone, or sit the full 60-question timed simulator.
Open the CCAO-F question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.