Home › Study guides › CCAR-P › Domain 4 › Lesson 4.5
CCAR-P · Domain 4 · 16% of the exam · Lesson 4.5 · 23 min read
Optimising tokens, latency and cost without losing quality
Where Claude's tokens and money go, how to split batch from interactive work, which levers cut cost or waiting, and how to prove each saving kept quality.
Written against objective 4.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.5.1 Why the bill tripled while the summaries stayed the same
Sablecrest is a media-monitoring company. Every day Claude summarises about 300,000 news articles for it, one short summary per article that mentions a client. Those summaries feed the client alerts: a bank learns that a regulator criticised it, a food brand that a recall rumour is spreading. Sablecrest also sells a research chat, where analysts ask questions such as "How did the regional press cover the launch this week?" and Claude searches the archive to answer.
In twelve months the Claude bill tripled, while article volume grew by about a third. Vesna, the architect Sablecrest brought in, found the reasons in the request logs. Every article went through the same synchronous pipeline as the chat, on the same model at its default settings. Each summary request carried the chat's twelve tool definitions, which the summariser never calls. The prompt asked for "a thorough summary" and got five paragraphs, after the model had thought at length about each article. Evander, the head of editorial, set one condition: whatever changed, the summaries clients receive must not get worse.
The instinctive fixes each pull one lever blindly: move everything to the cheapest model, cut the prompt, cap the output. Each can shrink the bill and quietly break the product. Cost-performance optimisation is the architect's alternative. You find where the tokens go and separate the work by who is waiting for it. You take the levers that cost no quality before the ones that trade it, and judge every change by what a USEFUL result costs, with the quality bar held.
4.5.2 Where the tokens and the money go
Claude bills per token, a piece of text of about three-quarters of an English word, with separate prices per million for input and output. It is tempting to read that price list as "input is cheap, so the answer is what we pay for". Half of that is right: on every current Claude model an output token costs five times an input token. But a request pays for everything it sends, every time it sends it, so input hides more than you expect.
| What you pay for | How it is billed | Where it hides |
|---|---|---|
| Model tier | Per million input / output tokens: $1 / $5 on Haiku 4.5, $2 / $10 on Sonnet 5.5, $4 / $20 on Opus 5.5, $10 / $50 on Fable 5.1 | Routine work sent to a tier above what it needs |
| Input | The model's input price | Instructions, examples, the document, and in a chat the whole conversation, resent on every turn |
| Tool definitions | Input: each tool's name, description and schema, plus a tool-use system prompt of a few hundred tokens | Tools attached to requests that never call them |
| Output | Five times the input price | Long free-text answers, and every retry of an output your code could not use |
| Thinking | Output: the reasoning Claude writes before answering, billed in full even when you display only a summary of it or nothing | An effort level higher than the work needs |
| Prompt cache | Writes cost 1.25 times input (five-minute lifetime) or 2 times (one hour); reads 0.1 times on most models, less on the newest top tiers | A prefix that changes on every request, so it is written and never read |
| Batch | Half price on input and output, and it stacks with caching | Work run synchronously that nobody is waiting for |
Memorise the shape: output is five times input, thinking is output, tool definitions are input on every request, batch halves everything and caching makes a repeated prefix cheap. Recognise the exact prices and cache multipliers; they vary by model and change with new releases.
Two of those rows are settings on the request. Prompt caching stores the processed form of a prompt's opening, the prefix every request repeats, so later requests read it at a fraction of the input price. Effort (low, medium, high, xhigh or max) tells Claude how many tokens to spend on thinking, tool calls and the answer itself; Claude Sonnet 5.5 defaults to high.
Vesna priced one average summary request on Claude Sonnet 5.5, at the list prices above. It sent about 8,300 input tokens: 2,300 of tool definitions and tool overhead, 4,500 of house rules and examples, and a 1,500-token article. It produced about 1,450 output tokens: 1,000 of thinking and a 450-token summary. That came to $0.031 per article, or $31.10 per thousand. A third of it paid for thinking, and about 15% for tools that were never called. Another 4% of summaries failed to parse and were sent again, paying the whole cost twice.
4.5.3 Split the work by who is waiting
Here is the question that decides more of the bill than any single setting: who is waiting for this output, and how long can they wait? Sablecrest ran one pipeline with one implicit answer, "right now", because the chat needed it. The summaries did not.
Vesna found three lanes. About 97% of summaries go into a client's morning digest, unread for hours. The other 3%, about 9,000 a day, match a client's crisis terms (a recall, a lawsuit, an executive's name) and must reach the client within minutes. And the analyst chat has a person at a screen.
The first lane is a textbook case for the Message Batches API: you submit up to 100,000 requests at once, they run asynchronously at half price, and most batches finish within an hour. Results arrive when the whole batch is done or after 24 hours, whichever comes first, and a request that has not run by then expires, unbilled. Think of the overnight post beside a courier: you pay courier rates only for what must arrive today.
Three lanes, decided by who is waiting
Digest summaries
late or expired: resubmit before the cut-off
Crisis alerts
Analyst chat
Two design details make the batch lane safe. Results come back in any order, so each request carries a custom_id that ties it to its article. And "most batches within an hour" is not "every batch", so Vesna submits a batch every hour and sets a 05:00 cut-off. At the cut-off her code resends any summary still missing or expired as an ordinary synchronous request, so the digest never waits on the slowest batch.
In the batch request, look at custom_id, at the one-hour cache_control on the shared house rules, and at output_config, which fixes the effort level and the JSON shape of every summary. The longer cache lifetime is the docs' advice for batches, whose requests run concurrently and get cache hits only on a best-effort basis.
batch = client.messages.batches.create(requests=[{
"custom_id": article.id, # results return in any order: match on this
"params": {
"model": "claude-sonnet-5-5",
"max_tokens": 4000,
"system": [{"type": "text", "text": HOUSE_RULES_AND_EXAMPLES, # identical in every request
"cache_control": {"type": "ephemeral", "ttl": "1h"}}], # one-hour cache for batches
"messages": [{"role": "user", "content": article.text}], # the only part that varies
"output_config": {
"effort": "medium", # chosen by a sweep on the eval set
"format": {"type": "json_schema", "schema": SUMMARY_SCHEMA}, # headline, bullets, sentiment
},
},
} for article in this_hours_articles])
4.5.4 Free wins before trade-offs
With the lanes split, which levers come first? Anthropic's cost guide sorts them into two groups. Free wins cut spend without touching quality: prompt caching, removing tokens that never influence the answer, batch for work that can wait, and auditing prompts written for an older model. Trade-offs exchange cost for intelligence: model choice, effort, output caps and multi-model designs. Take the free wins first. Your eval set, a fixed sample of real inputs with an agreed way to score each output, still checks them, but they need no quality argument.
| Lever | Cost | Waiting time | Quality risk |
|---|---|---|---|
| Prompt caching | A repeated prefix read at a tenth of the input price, or less | Faster first token on long prefixes | None |
| Remove unused tools and boilerplate | Fewer input tokens on every request | Slightly faster | None, if they never influenced the answer |
| Shorter, structured output | Fewer output tokens; no parse retries | Faster | Low: check the schema holds what readers need |
| Batch API | Half price | Up to 24 hours | None |
| Lower effort | Less thinking, fewer and terser tool calls | Faster | Real: sweep it on the eval set |
| Cheaper model or cascade | A fraction of the price on routine items | Faster; escalations wait twice | Real: needs a failure signal code can check |
| Streaming | Unchanged | First words appear sooner | None |
| Parallel tool calls | Lower when they replace round trips | The slowest call, not the sum | None, for independent read-only calls |
| Fast mode (some Opus models, research preview) | Premium price | Up to 2.5 times faster output | None |
Read it by column. Streaming buys time for nothing, fast mode buys time with money, and batch buys money with time. Caching, trimming and shorter output buy both, which is why they come first.
Vesna's first step took the free wins plus one low-risk change. She stripped the tools from summary requests, cached the house rules, opened the batch lane, and replaced five paragraphs with a JSON schema of headline, three bullets and sentiment. Structured outputs constrain the model's decoding to your JSON schema, so responses match it (refusals and truncated replies aside), and the 4% retry tax disappeared.
Then came the trade-offs, one at a time. An effort sweep, the eval set run at each effort level, showed that low cost quality and medium held it, and thinking fell from about 1,000 tokens a summary to about 400. Then a cascade: a cheap model handles every item first, and only the items that fail a check go to a stronger one. Picture a junior proofreader whose pages go through a checklist, with the senior editor seeing only the pages that fail it. The same pattern works inside one model: run everything at low effort and re-run only the failures at high.
The summary cascade
Routine articles
Checks fail
Known hard types
A cascade lives or dies on its escalation signal. Vesna's checks are ones code can run: the JSON validates, the summary names the client, and every organisation it mentions appears in the article. A model's rating of its own confidence is the model grading itself, not a check, and a check that passes bad work sends bad summaries out at the cheap price. Every escalated article pays twice and waits for a second batch, so the escalation rate is the cascade's cost driver. Where code can tell in advance that an article is hard, routing skips the cheap attempt: filings, long investigations and non-English text go straight to Sonnet 5.5.
4.5.5 Latency where a person waits
The chat is a different problem. It is about a tenth of the bill, and the complaint is time: at p95 (the wait that only one request in twenty exceeds), analysts waited 22 seconds for the first word. Vesna measured the wait as time to first token and total time per turn. Batch is useless here. A smaller model is the most direct latency lever, and Anthropic's latency guide lists it first, but it wins only where the task is narrow enough for the eval to hold. Analysts ask for judgement across many sources, and on the graded questions Haiku 4.5 fell short. So she looked for a lever at each stage of a turn.
Where an analyst's wait goes, and the lever for each part
medium held qualityStreaming sends the answer as it is generated, so the analyst reads while the rest arrives; it changes when the first words appear, not how many tokens you pay for. The chat's system prompt, tools and growing conversation are cached, which speeds up processing as well as cutting cost. Anthropic's effort guidance for chat and other latency-sensitive work on Sonnet 5.5 is to start at medium or low rather than the high default; medium held quality on 60 graded analyst questions.
Turns are the hidden multiplier. Every round trip resends the prompt, tools and conversation, and whatever Claude wrote earlier comes back as input on every later turn, so fewer turns cut cost and waiting at once. Claude can request several independent searches in one turn, and when your code runs them concurrently, the turn waits for the slowest search, not the sum. Lower effort helps here too, since it brings fewer and terser tool calls.
Finally, answers open with a brief of at most five sentences and an offer to expand; Anthropic's latency guide notes that asking for a number of sentences works better than a word limit. The result: p95 time to first token fell from 22 seconds to 5, and cost per session by about a third, with the graded answers unchanged. Fast mode was not the answer. It raises output speed on some Opus models at a premium price, in research preview, not the time to the first token that analysts complained about.
4.5.6 Cost per useful summary, not cost per token
The last trap is declaring victory on a smaller bill. A cheaper configuration that produces more bad summaries is not cheaper: a failed output still bills its tokens, then the retry, then an editor's time or a client's trust. So Vesna measures cost per useful summary: everything spent on the workload, including failed attempts, escalations and resubmissions, divided by the summaries that pass the quality bar. A factory judges a machine the same way, by its cost per part that passes inspection, scrap included. Here "useful" means the code checks pass and the summary meets the editors' rubric: faithful to the article, the client named correctly, the sentiment right.
Anthropic's cost guide gives the same rule for comparing models: compare cost per completed task, not per token. It adds that you should price the hardest tenth of your tasks, because that is where a cheaper model fails and the bill is decided. Per-token comparisons mislead in another way too: Claude models from 4.7 on turn the same text into about 30% more tokens than earlier ones such as Haiku 4.5. Here is Vesna's scorecard for Evander, every row measured on the same eval set.
Sablecrest digest summaries. Eval set: 400 articles weighted like real traffic. List prices, per 1,000 articles, every attempt included.
Baseline: Sonnet 5.5, synchronous, default effort, 12 chat tools attached, free-text summaries. Useful 91%. Parse retries 4%. $32.30 per 1,000 articles; $35.50 per 1,000 useful summaries.
Step 1, free wins plus a schema: tools removed, house rules cached for an hour, Batch API, JSON schema output. Useful 92%. Retries 0%. $9.90; $10.80 per 1,000 useful.
Step 2, effort: medium, after a sweep (low fell to 86% useful). Useful 92%. $6.90; $7.50 per 1,000 useful.
Step 3, cascade: Haiku 4.5 first; 15% of articles go to Sonnet 5.5, by code checks or article type. Useful 91%. $3.50; $3.85 per 1,000 useful.
Guardrail: useful rate at or above 90% on the eval set and on a weekly editor sample of 200 live summaries. Alert if escalations pass 25%.
Decision: ship steps 1 and 2 now. Run step 3 in shadow on 5% of traffic for two weeks, then decide.
Notice where the risk sits. Step 1 took most of the saving and traded nothing. Step 3 saved far less, and it depends on traffic the eval set may not contain, such as a new source with an unfamiliar layout. So it runs first in shadow, as Anthropic's measurement method recommends: on 5% of live traffic, scored but never sent to clients. The escalation alert matters for the same reason: if a new format starts failing the checks, the cascade can end up costing more than Sonnet alone.
4.5.7 The exam traps
Every trap here pulls one lever without asking what it moves, or reports a saving nobody measured.
- ✗ Turning on streaming to cut the bill. ✓ Streaming changes when the first words appear, not the tokens billed. For cost, reach for caching, batch, fewer tokens and effort.
- ✗ Sending work someone is waiting for through the Batch API. ✓ Batch only what can wait up to 24 hours. Keep users and deadlines of minutes on the synchronous path, and plan a cut-off for the batch tail.
- ✗ Moving everything to the cheapest model because its price per token is lowest. ✓ Compare cost per useful outcome on an eval set, hardest cases included. Cascade only behind a failure signal code can check.
- ✗ Lowering
max_tokensas the way to save money. ✓ The model cannot see the cap, so it does not write less; replies that hit it are cut off, still billed and redone. Ask for less (a schema, a shorter format, lower effort) and setmax_tokensas a ceiling with room for thinking and the full answer. - ✗ Deleting context the task needs to save input tokens. ✓ Remove what never influences the answer, such as unused tools, and cache what repeats. Required context stays.
- ✗ Reporting the lower bill as the result. ✓ Report cost per useful outcome against the baseline, with the quality guardrail that held.
4.5.8 Put it together: cut one workload's cost and prove quality held
You now have the whole method. Know where the tokens go, split lanes by who waits, take free wins before trade-offs, shorten the wait where a person waits, and let cost per useful outcome decide. The fastest way to own it is to run one workload through two configurations, then make the classic mistake yourself and watch the metric catch it.
This objective leans on its neighbours. Model selection (2.1) decides which tiers may enter a cascade, and context window management (2.4) keeps a long chat from resending what it no longer needs. Within this domain, monitoring (4.6) turns the scorecard's guardrail and escalation alert into live signals.
Key takeaways
- ✓ Output tokens cost five times input, thinking is billed as output, and every attached tool definition is input on every request, so find where the tokens go before choosing a lever.
- ✓ Split work by who waits: the Message Batches API halves the price of work that can wait up to 24 hours, while users and deadlines of minutes stay synchronous.
- ✓ Take the free and low-risk wins first (caching, removing unused tokens, batch, shorter structured output), then the trade-offs one measured step at a time (effort, routing, a cheaper model or a cascade).
- ✓ A cascade needs a failure signal that code can check, and its escalation rate drives its cost.
- ✓ Where a person waits, streaming, a cached prefix, matched effort, fewer turns and shorter answers cut the wait, though streaming on its own lowers no cost.
- ✓ Judge every change by cost per useful outcome against a baseline on the same eval set, with quality held as a guardrail and watched after launch.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
30 CCAR-P questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 30 questions on Evaluation, Testing & Optimization alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.