Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 5 › Lesson 5.4

CCDV-F · Domain 5 · 16.8% of the exam · Lesson 5.4 · 22 min read

Cost and token management: model the bill, then cut it

Why a Claude bill outgrows its forecast, and how to budget tokens, track usage, model monthly cost and cut it with cache checkpoints and batches.

Written against skill 5.4 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.4.1 Why the bill tripled when nothing broke

Kestrel Broadband, a regional internet provider, launched a support chatbot built on Claude. It answers billing questions and walks customers through router problems, and it works: fewer calls reach the phone team, and customers rate the answers well. Then the first full month's invoice arrives. Graciela, who owns the support budget, had set aside about $2,000 a month, based on the pilot. The bill says about $6,000. Traffic is exactly what was forecast. So where did the money go?

Nothing broke. The bill grew because of how a chat application pays. The Claude API charges per token, the small chunk of text a model reads and writes (in English, roughly four characters or three quarters of a word). Input and output tokens are priced separately.

And the API is stateless: it remembers nothing between requests. So every turn of a conversation sends everything again: Kestrel's 6,000-token system prompt (the standing instructions and policies at the top of every request), every earlier message and reply, then the new question. An eight-turn chat pays for that system prompt eight times. On top of that, every night the team replays hundreds of recorded conversations to check quality, at full price, and nobody budgeted for those runs.

Anouk, the developer who built the chatbot, does not need a cheaper model first. She needs to know where the tokens go. Cost and token management is that discipline. You budget what each request may use, track what every request really used and turn the numbers into a cost model. Then you pull the levers that shrink the biggest lines, chiefly prompt caching and batch processing.

What the budget counted, and what the month billed

The budget counted

Short pilot chatsfour turns each
Question and answernothing else
About $2,000 a month

The month billed

The 6,000-token system promptresent on every turn
The whole historyresent, and growing
Nightly evaluation runsat full price
About $6,000 a month
The budget priced short pilot chats; the real bill pays for the system prompt on every turn, the growing history and the nightly test runs.

5.4.2 Budget the tokens before you send

Here is the question most first applications never ask: how big is this request, and what may it cost, before it goes out? A token budget is the limit you set on that, per request, per conversation or per feature, and it has two sides.

The input side is what you send: system prompt, history, tools, the new message. You can measure it in advance with the token counting endpoint. It takes the same system, messages and tools as a real request and returns input_tokens. It is free to use, with its own rate limits, and its figure is an estimate. It counts for the model you name, which matters. Claude 4.7 and later models use a newer tokenizer, the part that splits text into tokens, and it produces about 30% more tokens for the same text than earlier models. Recount whenever you switch models.

Kestrel needs this because customers paste router logs, and some logs run to tens of thousands of tokens. Look at the last three comments: the count happens before the paid call, an oversized message is trimmed first, and max_tokens sits on the reply. keep_error_lines stands for your own trimming code.

BUDGET_IN = 20_000                        # context budget for one chat turn

def send_turn(history, message):
    request = dict(model=MODEL, system=PLAYBOOK,
                   messages=history + [{"role": "user", "content": message}])
    size = client.messages.count_tokens(**request).input_tokens  # free estimate
    if size > BUDGET_IN:                                         # CHECK before you pay
        message = keep_error_lines(message)
        request["messages"] = history + [{"role": "user", "content": message}]
    return client.messages.create(**request,
                                  max_tokens=1024)               # a cap on the reply, not a saving

The output side is max_tokens, the most tokens one reply may contain. It is tempting to lower it to save money. Resist that. The model never sees the cap and does not write more briefly because of it. The reply stops mid-sentence, the response's stop_reason field reads "max_tokens", you pay for the tokens it did write, and the customer asks again. Think of max_tokens as a fuse, not a thermostat: it stops a runaway reply, but it never lowers the everyday bill. Shorter replies come from asking for them in the prompt.

5.4.3 Track what every request really used

The invoice says what the month cost. It cannot say which feature spent it or whether last week's change helped. For that you need a record per request, and the API gives you one: every response carries a usage object with the exact token counts billed for that call. The token count you took beforehand was an estimate; usage is the truth.

Two of its four fields belong to prompt caching, a discount on the opening part of a prompt when you resend it unchanged; a later section shows how to set it up. For now, read a cache breakpoint as the marker where that reusable part ends.

Field in usage What it counts How it is priced
input_tokens Input not read from or written to the cache: with caching on, only what comes after the last cache breakpoint The base input price
cache_creation_input_tokens Input written to the cache on this call, split by lifetime in cache_creation 1.25x the input price (5-minute) or 2x (1-hour)
cache_read_input_tokens Input read back from the cache 0.1x the input price on most models
output_tokens Everything generated, including any thinking The output price, five times the input price on current models

Memorise the first row's trap. Once caching is on, input_tokens is no longer your total input. Total input is input_tokens plus both cache fields. A dashboard that reads only input_tokens will show input spend collapsing the day caching goes live, while cache writes and reads are still on the bill.

Anouk logs usage for every request with a tag for the feature, chat or eval. For the totals, the Claude Console has Usage and Cost pages. The Usage and Cost API returns the same data to your own scripts, which can group it by model, workspace or API key; data typically appears within about five minutes. It needs an Admin API key, an organisation-level key separate from the one the chatbot uses. A separate API key for the evaluation runs makes the split visible there too.

One more control belongs to the budget. A workspace, a group of API keys inside your organisation, can carry a spend limit. Once spending reaches it, the API refuses further requests until the limit resets or someone raises it. That is the backstop: it protects the budget, but it saves nothing.

5.4.4 Model the cost: request, conversation, month

With a day of usage logged, Anouk can answer the question Graciela actually asked: why $6,000? A cost model turns token counts into money at three levels, each built from the one before.

Per request, you price each usage count at its own rate. Anouk uses the list price of Claude Sonnet 5.5, the model Kestrel runs on, at the time of writing: $2 per million input tokens and $10 per million output tokens. Look at the four lines inside the sum: each count gets its own multiplier, and the cache fields are priced from the input rate.

PRICE_IN, PRICE_OUT = 2.00, 10.00            # $ per million tokens; read them from the pricing page
READ, WRITE_5M, WRITE_1H = 0.1, 1.25, 2.0    # cache multipliers on the input price

def request_cost(u) -> float:
    writes = u.cache_creation                 # cache writes, split by lifetime
    w5 = writes.ephemeral_5m_input_tokens if writes else 0
    w1 = writes.ephemeral_1h_input_tokens if writes else 0
    total = (u.input_tokens * PRICE_IN                          # after the last breakpoint
             + (u.cache_read_input_tokens or 0) * PRICE_IN * READ
             + (w5 * WRITE_5M + w1 * WRITE_1H) * PRICE_IN
             + u.output_tokens * PRICE_OUT)                     # replies, and any thinking
    return total / 1_000_000

record(feature="chat", usd=request_cost(response.usage))       # your metrics call: tag every request

Per conversation, the stateless API does the damage. The average Kestrel chat has eight turns; a customer message is about 150 tokens and a reply about 300. Turn 1 sends 6,150 tokens. Turn 8 sends the 6,000-token prompt, seven earlier exchanges (3,150 tokens) and the new message: 9,300 tokens. Over the whole chat that is 61,800 input tokens, of which 48,000 are the same system prompt, plus 2,400 output tokens: about 15 cents. The history part grows faster than the number of turns, because each new turn resends every turn before it.

Per month, you multiply by volume and add the scheduled jobs. Kestrel handles 30,000 conversations a month. Every night the evaluation suite replays 500 recorded six-turn conversations and asks for each of their six replies again, 3,000 requests a night.

Cost driver Tokens a month Cost a month
System prompt, resent on every turn 1.44 billion input $2,880
Earlier turns, resent as history 378 million input $756
New customer messages 36 million input $72
Replies 72 million output $720
Nightly evaluation runs 655 million input, 27 million output $1,580
Total about $6,000

Read the table from the top. Almost half the bill is one block of text that never changes. A quarter is work nobody watches. The replies cost ten times as much as the customer messages: twice the tokens at five times the price. Each line points at a different lever.

5.4.5 Cache the prefix and place the checkpoints

The biggest line is 6,000 identical tokens, processed at full price 240,000 times a month (30,000 chats of eight turns). Prompt caching stops that. The API keeps its processed form of the beginning of a prompt, the prefix, and reuses it when a later request starts with exactly the same content, in the order tools, system, messages. You pay a premium once to write the entry, then a fraction on every read. An entry lives for five minutes by default, or for an hour if you pay more to write it; that lifetime is its TTL (time to live).

Cache operation Price, relative to base input When it pays
Write, 5-minute lifetime 1.25x After one read
Write, 1-hour lifetime 2x After two reads
Read (a cache hit) 0.1x on most models; a few newer models read for less Every reuse, and each read restarts the TTL at no charge

Memorise the three multipliers; the cheaper reads on a few newer models you only need to recognise. Cache what is stable and reused: tool definitions, system instructions, policy text, examples. Anything that changes per request, such as the customer's name, the date or an account number, goes after the shared part, in the newest user message. One timestamp at the top of the system prompt makes every prefix unique, so every request pays the write premium and never reads: Anthropic measured a run like that costing more than no caching at all.

You mark where a cached prefix ends by adding cache_control to a content block, one piece of the request such as the system text or a single message. The docs call that mark a cache breakpoint; the exam guide's term for placing them is cache check-pointing. A request can carry up to four.

The rule that decides where they go: a write happens only at a breakpoint, and a later request can read only an entry an earlier request wrote. Think of save points in a video game: you can restart only from a spot where someone actually saved, never from a spot you merely walked past. So each checkpoint belongs at the end of a section that stays identical across the requests meant to share it.

Kestrel needs two. Checkpoint 1 sits on the system prompt, so every customer's chat can read the one entry the others keep warm. Checkpoint 2 follows the conversation: automatic caching, a single cache_control at the top level of the request, marks the last block and moves forward each turn. With only the automatic one, each new chat would write the 6,000 tokens again, because no request ever wrote an entry ending at the system prompt.

Two checkpoints in one chat request

Checkpoint 1: every customer

The 6,000-token system promptidentical for all
Read at 0.1xkept warm by other chats

Checkpoint 2: this conversation

Earlier turnsread at 0.1x
Last reply and new messagewritten at 1.25x

moves to the newest block every turn

The system prompt is cached once for every customer; each conversation's history is cached as it grows, so a turn writes only its newest blocks.

In code, look at the two cache_control entries: one on the system block, one at the top level for the conversation. The last line is how you verify hits instead of hoping for them.

response = client.messages.create(
    model=MODEL,
    max_tokens=1024,
    cache_control={"type": "ephemeral"},          # checkpoint 2: moves with the conversation
    system=[{
        "type": "text",
        "text": PLAYBOOK,                         # 6,000 tokens, the same for every customer
        "cache_control": {"type": "ephemeral"},   # checkpoint 1: shared by every chat
    }],
    messages=history + [{"role": "user", "content": customer_message}],
)
u = response.usage
print(u.cache_read_input_tokens, u.cache_creation_input_tokens, u.input_tokens)

The result: a chat now reads 58,500 of its 61,800 input tokens at 0.1x and writes 3,300 at 1.25x. It costs about 4.4 cents instead of 15, and the four chat lines of the monthly table fall from about $4,430 to about $1,320.

The TTL is the last decision. Customers told to restart their router often come back after ten minutes, when their conversation's entry has expired, so that turn pays the write price for the whole history again. Anthropic's cost guide suggests the 1-hour TTL once more than about one gap in twenty between turns falls between five minutes and an hour, as long as gaps over an hour are rare. Anouk counts those gaps in a day of logged traffic before she switches.

5.4.6 Batch what can wait, and the other levers

The second-largest line, the nightly evaluation runs, has a property the chat does not: nobody is waiting for it. The Message Batches API fits that shape. You submit many Messages requests as one job and Anthropic works through them in the background. Most batches finish within an hour, and any request not processed within 24 hours expires. Every token costs 50% of the standard price, input and output alike, and the discount stacks with caching. Moving the evaluation runs to a batch halves that line from about $1,580 to $790. The chat stays on the realtime API, because a customer is watching the screen.

With caching and batching, Kestrel's month drops from about $6,000 to about $2,100, near the budget, with no change in answer quality. The remaining levers each target a different driver, and not all are free.

Cost driver Lever What it costs you
The same prefix on every request Prompt caching Nothing in quality; needs a stable prefix
Work nobody waits for Message Batches API, 50% off Results arrive within 24 hours, not seconds
Long replies Ask for shorter answers in the prompt Test that the short answer still resolves the issue
A bloated system prompt Audit it and cut what the current model no longer needs A before-and-after test run
History that keeps growing Prune or compact old turns Some context is summarised or dropped
Simple routes on a large model A smaller model for those routes Quality must be tested per route

Shorter answers pay twice in a chat. An output token costs five times an input token, and every reply is resent as input on each later turn. Anthropic measured a one-line answer format against a five-section memo on the same task: the memo cost 2.8 times as much, for accuracy within noise.

One tempting lever is missing from the table on purpose. Streaming, which sends the reply to the screen piece by piece as it is generated, makes the chatbot feel faster but bills exactly the same tokens. It is a choice about user experience, not a saving.

5.4.7 The exam traps

Every trap here either guesses instead of measuring, or pulls a lever that misses the real driver.

  • ✗ Estimating cost from one question and one answer. ✓ Model whole conversations from logged usage. Every turn resends the system prompt and the history, and scheduled jobs such as evaluations belong in the monthly total.
  • ✗ Lowering max_tokens to cut spend. ✓ Keep max_tokens as a safety cap sized for a complete answer. The model does not see it; a truncated reply is still billed and usually asked for again.
  • ✗ Reading input_tokens as total input once caching is on. ✓ Add cache_creation_input_tokens and cache_read_input_tokens, and price each at its own rate.
  • ✗ Putting per-request details at the top of the cached prompt. ✓ Stable content before the checkpoint, variable content after it. A changing line at the top turns every request into a cache write.
  • ✗ Sending live chat through the Message Batches API for the discount. ✓ Batch only work nobody is waiting for; a batch may take up to 24 hours.
  • ✗ Switching everything to the smallest model first. ✓ Take the free wins (caching, batching, trimming) first, and choose a model per route from tested requirements.

Four tempting cuts, one real plan

Lower max_tokenstruncated, still billed
Stream the repliesthe same tokens billed
Smallest model everywherequality falls first
Batch the live chatcustomers wait hours
Measure, then cut the biggest linecache what repeats, batch what can wait
Each tempting cut changes something beside the driver; the real plan measures usage first and pulls the lever built for the biggest line.

5.4.8 Put it together: model a bill and cut it

You now have every piece: a budget for each request, a usage record for every call, a cost model from request to month, and the levers matched to the lines they shrink. To make it stick, measure a small chat, break its cache on purpose, and watch the cost jump.

Two neighbouring skills own the levers this lesson only names. Model selection and trade-offs (5.3) picks a smaller model for a simple route by requirement, not by price alone. Context engineering (6.1) prunes and compacts a long history so conversations stop growing without limit. Both show up in the same usage numbers you now log.

Key takeaways

  • ✓ You pay per token, input and output priced separately, and every turn of a conversation resends the system prompt and the whole history.
  • ✓ A token budget limits each request: count input with the free token counting endpoint before sending, and treat max_tokens as a safety cap, not a saving.
  • ✓ The usage object is the truth for cost; once caching is on, total input is input_tokens plus cache writes plus cache reads, each priced differently.
  • ✓ Model cost per request, per conversation and per month from measured usage, including scheduled jobs; the largest line picks the first lever.
  • ✓ Prompt caching writes a stable prefix at 1.25x (5-minute TTL) or 2x (1-hour TTL) and reads it at 0.1x on most models, with up to four checkpoints placed where content stops changing.
  • ✓ The Message Batches API halves the price of work nobody waits for and stacks with caching; shorter outputs, trimmed prompts and smaller models come next, each measured.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

27 CCDV-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 27 questions on Model Selection and Optimization alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources