Home › Study guides › CCDV-F › Domain 2 › Lesson 2.3
CCDV-F · Domain 2 · 33.1% of the exam · Lesson 2.3 · 24 min read
Claude API mechanics: requests, streams, caches and batches
How a Messages API call works, and when to stream it, add images or thinking, cache a prefix, batch it, or send it through Bedrock, Vertex AI or Foundry.
Written against skill 2.3 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
2.3.1 Why "just call Claude" is not a design
Ifeoma is the only back-end developer at a news-summary startup. Readers ask the app about the day's news and get short answers in the house voice. A year in, the product makes five different demands on the API. A chat answers readers live, and an upload box takes screenshots of articles. A nightly job writes 20,000 personalised digests before breakfast. A house style guide of about 9,000 tokens (the word pieces a model reads, writes and bills by) rides along on every call. And one enterprise client's contract says its requests must run through Amazon Bedrock, Amazon's hosted AI service, not through Anthropic directly.
Ifeoma's first version had one helper, ask_claude(text), and every feature called it the same way. Readers stared at an empty chat bubble for fifteen seconds while long answers finished. The nightly job ran for hours at full price, and the style guide was billed in full on every call. Then the client's security team asked which of these features even exist on Bedrock, and nobody knew.
None of this was a model problem; the integration around the model was wrong. Every feature uses the same thing, the Messages API (API: application programming interface), a single request that carries a conversation and returns Claude's reply. What changes is HOW you send it, and each way of sending it answers one requirement.
One request, five ways to send it
2.3.2 What a request carries and what comes back
Here is the belief behind many teams' first bug: Claude remembers the conversation. It does not. The Messages API is stateless: each request is complete in itself, and the service keeps nothing between calls. If a reader asks "and what did the minister say?", your code must resend the earlier turns, or Claude has no idea which minister. Think of a translator hired by the page who forgets you the moment each page is done: every job arrives with its whole brief.
A request has four parts you set on almost every call. The model field names the Claude model. The max_tokens field is the hard ceiling on how many tokens Claude may write; it can stop sooner, never later. The messages list is the conversation so far, turns with the role user or assistant, ending with the new user turn. Standing instructions such as the style guide go in the top-level system field. Content can be a plain string or a list of content blocks: text, images, documents and tool results.
The reply mirrors that shape: a list of content blocks, plus two fields your code must read before it trusts the answer.
| Reply field | What it tells you | What Ifeoma's code does with it |
|---|---|---|
content |
The blocks Claude wrote: text, tool_use, thinking |
Shows the text; hands tool_use blocks to the code that runs tools |
stop_reason |
Why Claude stopped: end_turn, max_tokens, tool_use, stop_sequence, pause_turn, refusal, model_context_window_exceeded |
Accepts end_turn; treats a digest that ended on max_tokens as cut off, not finished |
usage |
input_tokens, output_tokens and the cache counters |
Logs them per feature, so every cost has a source |
Memorise the three fields and what end_turn, max_tokens and tool_use mean; recognise the other stop reasons when you see them.
Tools travel through the same exchange. You pass a tools list, each tool with a name, a description and an input_schema. When Claude wants one, the reply ends with stop_reason: "tool_use" and holds a tool_use block naming the tool and its input. Your code runs the real function and returns the answer in the next user turn as a tool_result block whose tool_use_id quotes that block's id. Server tools such as web search are the exception: Anthropic runs them for you.
2.3.3 Streaming and thinking: the answer as it forms
Now the frozen chat. Without streaming, your code waits for the last token and returns everything at once, so the reader watches an empty bubble. With streaming, you set "stream": true or use the SDK's messages.stream() helper (SDK: software development kit). The API then sends the reply while it is written, as server-sent events (SSE), a standard way for a server to push small messages down one open HTTP connection.
The events always arrive in the same order, shown below. One detail needs care in code: an error such as overloaded_error can arrive mid-stream, after the HTTP status already said 200. Your code must handle errors inside the stream, not only before it starts.
The event order in one streamed reply
message_startthe reply, content still emptycontent_block_starta text, tool_use or thinking block beginscontent_block_deltanew content, many times, then content_block_stopmessage_deltastop_reason and cumulative usagemessage_stopthe stream endsStreaming is a latency tool, not a cost tool: the same tokens are generated and billed, but the reader sees the first words sooner. It is also the safe way to request long outputs. Anthropic's docs recommend streaming or batching for long requests, and the SDKs check that a non-streaming request is not expected to run past ten minutes. One rule for tools: a tool_use block's input streams as fragments of JSON, so act on it only once the block is complete.
In the chat handler, look at text_stream, which yields each piece of text as it arrives, and get_final_message(), which returns the complete reply once the stream ends.
with client.messages.stream(
model=MODEL,
max_tokens=2048,
system=STYLE_GUIDE,
messages=history, # the full conversation, as always
) as stream:
for text in stream.text_stream: # each fragment as Claude writes it
send_to_browser(text) # the reader sees words straight away
final = stream.get_final_message() # the same Message that create() returns
history.append({"role": "assistant", "content": final.content})
log_usage(final.usage) # same tokens, same bill as without streaming
if final.stop_reason == "max_tokens":
send_to_browser("\n[answer cut short]")
The same stream can carry a second kind of block. With thinking, Claude reasons through a hard question in thinking blocks before it writes the answer. That helps a reader asking how a rate decision will ripple through the housing market, and is wasted on "summarise this headline". Newer models use adaptive thinking, thinking: {"type": "adaptive"}: Claude decides whether and how deeply to think, guided by the effort setting (how hard you ask Claude to work). On the newest models it is on by default. The Claude 4.5 generation and earlier use extended thinking, {"type": "enabled", "budget_tokens": N}, a budget you set of at least 1,024 tokens and below max_tokens. Claude 4.7 and later models reject that form.
Four mechanics matter at the API level. Thinking tokens are billed as output and count toward max_tokens, so leave room for the answer too. A thinking block holds a summary of the reasoning, or an empty field when display is "omitted", the default on newer models; you pay for the full thinking either way. Inside a tool-use exchange, you pass thinking blocks back exactly as you received them. And in a stream, the thinking block comes first, as thinking_delta events, before any text.
2.3.4 Images and documents: getting data in front of Claude
Readers upload a screenshot of an article and ask for the gist. A picture enters a JSON request as a content block: an image block in the user turn, whose source says where the pixels come from. That can be base64 data (the file's bytes encoded as text) embedded in the request, or a url. It can also be a file_id from the Files API, where you upload a file once and reference it by ID afterwards. Claude reads JPEG, PNG, GIF and WebP, and works best when the image comes before the question.
Screenshots bring two traps. The first is size: an image costs tokens roughly in proportion to its area, and one larger than the model's maximum resolution is scaled down first. A full-page capture is expensive and, once scaled, its text may be too small to read, so crop to the article. The second is repetition: because the API is stateless, a base64 image is resent in full with every follow-up question. Uploading it once and sending the file_id keeps each request small.
That generalises into the data access patterns of the Messages API: the ways your application puts information in front of Claude.
| Pattern | How it works | Use it when |
|---|---|---|
| Inline | Text in the message, or base64 image or PDF data in a content block | The content is small or used once |
| By reference | A url source, or a file_id from the Files API |
The same file comes back across turns or requests |
| On demand | A tool your code runs returns the data in a tool_result |
The data is large, private or changing, and Claude should fetch only what it needs |
| Grounded | document or search_result blocks with citations enabled |
Each claim must point back to the exact source passage |
Learn when each pattern fits; recognise the block names.
Ifeoma uses the last row for the digests. Each source article goes in as a document block with "citations": {"enabled": true}, and the reply's text blocks carry citations pointing at the exact sentences used, so an editor can check any claim. When the chat's own search tool returns articles as search_result blocks, each with a source, a title and text, Claude cites them the same way.
The Files API adds one rule. Uploaded files belong to your whole workspace (the group your API keys belong to), not to one reader, so any key in that workspace can use any file_id. If the browser could send a file_id, one reader could summarise another reader's upload. Keep the mapping from readers to files in your own database, and never accept a file_id from the client.
2.3.5 Prompt caching: pay once for the part that never changes
The 9,000-token style guide goes out on every chat turn and every digest, identical each time, and Ifeoma pays full price for the API to process it again and again. Prompt caching lets the API reuse its processed form of a prompt's beginning, its prefix, across requests. The first request writes the prefix to the cache. Later requests that start with exactly the same prefix read it back, faster and at a fraction of the input price.
Think of a print shop that keeps a ready-made plate for your letterhead. Every letter that starts with that exact letterhead skips the typesetting, but print the date above it and each letter needs a new plate.
Three rules make it work. First, the prefix is built in a fixed order, tools, then system, then messages, and a change at any point invalidates the cache from there on. Second, you mark where the reusable part ends with a cache breakpoint, "cache_control": {"type": "ephemeral"} on a content block, and everything up to it must match exactly. You can set up to four breakpoints, or put one cache_control at the top level of the request and let the API place it on the last cacheable block. Third, an entry lives five minutes by default and each hit refreshes it at no cost; a one-hour lifetime costs more to write.
What gets cached, and what must come after
Cached prefix
toolstool definitionssystemthe 9,000-token style guidecache_control on the last stable blockAfter the breakpoint
Ifeoma's first attempt got no hits. She had put "Today is 29 September. Reader ID 48213." at the top of the system prompt, above the style guide. That line changed on every request, so the prefix never matched, and every call paid for a fresh cache write, which costs a little more than not caching. The fix was ordering. In the version below, the breakpoint sits on the style guide and the date comes after it.
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system=[
{"type": "text", "text": STYLE_GUIDE, # identical on every call
"cache_control": {"type": "ephemeral"}}, # breakpoint: cache up to here
{"type": "text", "text": f"Today is {today}."}, # changes daily, so it goes AFTER
],
messages=history,
)
u = response.usage
print(u.cache_creation_input_tokens, # tokens written to the cache on this call
u.cache_read_input_tokens, # tokens reused from it: the cheap part
u.input_tokens) # tokens after the breakpoint, full price
The usage counters are how you verify caching rather than hope for it. If both cache counters stay at zero, nothing was cached. Often the prefix is shorter than the model's minimum cacheable length (512 to 4,096 tokens, depending on the model), and the API then skips caching without an error. Caching never changes the answer, only the cost and speed of reading the prompt.
2.3.6 Realtime, batch, or another cloud
The nightly digests are the chat's opposite: nobody watches them being written, there are 20,000 of them, and they are due hours away, at breakfast. The Message Batches API fits that shape. You submit many Messages requests in one call, they are processed asynchronously, and you collect the results later. Each request carries your own custom_id and ordinary Messages params. A batch holds up to 100,000 requests or 256 MB, most finish within an hour, and anything unfinished after 24 hours expires. Everything in a batch costs half the standard price.
Batching changes how you handle results. They come back in any order, so you match each one to its request by custom_id, never by position. Each request succeeds or fails on its own, as succeeded, errored, canceled or expired, and the last three are not billed. An expired request or a server error can go again as it is; an invalid one must be fixed first. The batch outlives the process that created it, so you store its ID and poll processing_status until it reads ended. In the code, look at the custom_id on the way in and on the way out.
batch = client.messages.batches.create(requests=[
{"custom_id": f"digest-{r.id}", # your key for matching later
"params": {"model": MODEL, "max_tokens": 800,
"system": STYLE_SYSTEM, # cached style guide, same in every request
"messages": [{"role": "user", "content": r.brief}]}}
for r in readers
])
save_batch_id(batch.id) # the job outlives this process
# Later, from a scheduled job:
if client.messages.batches.retrieve(batch_id).processing_status == "ended":
for res in client.messages.batches.results(batch_id): # any order
if res.result.type == "succeeded":
store_digest(res.custom_id, res.result.message) # match by ID, not position
else:
retry_later(res.custom_id, res.result) # fix if invalid, resubmit only this one
| Messages API (realtime) | Message Batches API | |
|---|---|---|
| Answer arrives | In seconds, streamed if you like | Asynchronously: most within an hour, at most 24 hours |
| Price | Standard | Half the standard price |
| Unit of failure | The call: retry it | Each request: succeeded, errored, canceled or expired |
| Matching | The reply belongs to its call | By custom_id; order is not kept |
| Best for | A person waiting | Bulk work whose deadline is hours away |
Batching trades time for price. With a hard deadline inside the 24-hour window, submit early and keep a fallback for requests that expire. With a person waiting, or a deadline shorter than a batch can promise, stay realtime.
The last requirement comes from the enterprise client. Claude is also offered through three cloud platforms: Amazon Bedrock, Google Cloud (Vertex AI, which Anthropic's docs now call Agent Platform) and Microsoft Foundry. The request body is essentially the Messages API you know, and Anthropic's Python and TypeScript SDKs have a client for each. What changes is identity, naming and billing. You authenticate with AWS (Amazon Web Services) credentials, Google Cloud credentials, or an Azure key or Microsoft Entra ID token, instead of an Anthropic API key. Model names take an anthropic. prefix on Bedrock, sit in the endpoint URL on Vertex AI and become your deployment name on Foundry. The bill arrives through the cloud account.
Features differ too, which is what the security team was really asking. The Message Batches API is not available on any of the three. The Files API and URL image sources are not available on Bedrock or Vertex AI, so images go inline as base64; on Foundry, the Files API works only on deployments hosted by Anthropic. Ifeoma routes the client's chat through a Bedrock client with base64 screenshots, keeping streaming and caching, and runs its digests as paced realtime calls through the night.
Same request, different front door
Stays the same
Changes
Not available
2.3.7 The exam traps
Almost every mistake in this skill applies a real mechanism to the wrong requirement.
- ✗ Streaming to reduce cost. ✓ Stream to cut the wait for the first word. The same tokens are billed; cost falls by caching repeated input or batching work that can wait.
- ✗ Putting per-request details in front of the cached content. ✓ Stable content first, the breakpoint on the last block that never changes, everything variable after it. A date at the top makes every prefix unique.
- ✗ Sending a waiting user's request through the Batches API because it is cheaper. ✓ Keep people on the realtime Messages API. A batch may take up to 24 hours.
- ✗ Matching batch results by position, or rerunning the whole batch when a few fail. ✓ Match by
custom_id, and resubmit only the requests that failed: fix the invalid ones, resend the expired ones. - ✗ Assuming a cloud vendor is the Claude API under another name. ✓ The request shape is shared, but check the provider's feature list: no Batches API on Bedrock, Vertex AI or Foundry, and base64-only images on Bedrock and Vertex AI.
- ✗ Treating a reply as finished without reading
stop_reason. ✓ Amax_tokensstop means cut off, and with thinking on, the reasoning spends from that same budget.
2.3.8 Put it together: send one request five ways
You now have every piece. The quickest way to make them stick is to send one request through each mechanism, watch the numbers in usage, and break the one that fails most quietly.
Everything else in this domain builds on these calls. Software engineering foundations (2.4) wraps them in asynchronous code, retries and version control. Claude application design (2.5) decides what belongs in system and what in each turn, and keeps one reader's context out of another's session. Configuration management (2.6) pins the model versions these requests name. Beyond the domain, tool implementation (8.1) builds the loop around tool_use, and cost and token management (5.4) does the arithmetic on caching and batching.
Key takeaways
- ✓ The Messages API is stateless: each request carries the model,
max_tokens,systemand the full history, and each reply returns content blocks,stop_reasonandusage. - ✓ Tools use the same exchange: Claude returns a
tool_useblock, your code runs the tool and sends back atool_result. - ✓ Streaming delivers the reply as server-sent events, cutting the wait for the first word but not the cost; thinking adds
thinkingblocks whose tokens are billed as output and count towardmax_tokens. - ✓ Images and documents are content blocks sourced inline, by URL or by Files API
file_id; tools fetch data on demand, and citations tie answers to exact passages. - ✓ Prompt caching reuses an exact prefix (
tools,system,messages) up to acache_controlbreakpoint, so stable content goes first and hits show incache_read_input_tokens. - ✓ The Message Batches API halves the price of work that can wait up to 24 hours; match results by
custom_idand resubmit only the failures. - ✓ Bedrock, Vertex AI and Foundry keep the request shape but change authentication, model naming and billing, and lack some features, including the Batches API.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
51 CCDV-F questions on Domain 2, free
Every question in the bank is tagged to a domain, so you can drill 51 questions on Applications and Integration alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 2 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.