Home › Study guides › CCAR-F › Domain 4 › Lesson 4.5
CCAR-F · Domain 4 · 20% of the exam · Lesson 4.5 · 20 min read
Batch processing with the Message Batches API
When the Batches API's 50% saving is worth a wait of up to 24 hours, how custom_id matches results, and how to plan an SLA and resubmit only what failed.
Written against task statement 4.5 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.5.1 Who is waiting for the answer?
Picture a software team that has built Claude into its continuous integration (CI) pipeline: the automated system that checks every code change before it joins the main code. Claude has two jobs there. The first is a pre-merge check. Claude reviews each pull request (a proposed change waiting for approval), and the change can't be merged until the review passes. A developer sits there waiting.
The second is a nightly test-gap report. At midnight a job sends each module of the code (a self-contained part, such as payments or login) to Claude. Claude lists the behaviour no test covers and drafts the missing test cases. The QA lead reads the report over breakfast. Nobody waits on it while it runs; everybody is asleep.
Both jobs make ordinary real-time calls today: send a request, keep the line open, get the answer in seconds. Then the quarterly API bill arrives, and the engineer who runs the pipeline notices that Anthropic's Message Batches API handles the same requests at half the price. Her proposal: move both jobs over. Same model, same prompts, half the bill. What could go wrong?
Think about what a real-time call is paying for. Part of the price buys the promise of an answer NOW. For the pre-merge check, that promise is the whole point. For the test-gap report it's wasted: the answer arrives in seconds and then sits unread for eight hours. The Message Batches API drops that promise. You hand over a large pile of requests in one go, and Anthropic works through them in the background at half the price, taking up to 24 hours.
4.5.2 What you give up for half the price
Here is the thought that tempts people, and the exam builds questions around it: most batches finish within an hour, so why not batch everything? To see why not, compare how the two APIs behave.
With the ordinary Messages API, your code sends one request and waits on the same connection until Claude's reply comes back. Engineers call this synchronous. With the Message Batches API, your code sends a whole list of requests, and what comes back straight away is not an answer but a receipt: a batch ID and a status of in_progress. Anthropic then processes each request independently in the background (engineers say asynchronously), and you come back later for the results. The documented ceiling is 24 hours; a request that still hasn't run when the 24 hours are up expires. There is no latency (waiting time) service-level agreement (SLA): nothing promises an answer by any earlier time.
Think of a print shop with two counters. At the express counter you wait while they print your handout, and you pay full price. At the overnight counter you drop off a box of jobs, collect a ticket and pay half. The jobs are usually ready within the hour, but the shop only promises "within a day". Nobody takes the handout for a meeting in twenty minutes to the overnight counter, however cheap it is.
| Batches API fact | What it means in plain terms | What it rules in or out |
|---|---|---|
| 50% cost saving | Every input and output token (the unit of text the API bills by) costs half the standard price | Worth it for large volumes of work that can wait |
| Up to 24-hour processing window | Most batches finish within an hour; the ceiling is 24 hours | Plan deadlines on 24 hours, never on the typical time |
| No guaranteed latency SLA | Nothing promises a result by any earlier time | Rules out anything a person or a pipeline blocks on |
custom_id on every request |
Your own label, returned with each result, because results can arrive in any order | Match results to requests by custom_id, never by position |
| No multi-turn tool calling | Your code can't run a tool mid-request and hand back the result | Rules out agents that call your tools as they work |
All five rows are exam facts. The first three decide WHETHER to batch; the last two shape HOW you build a batch.
The last row is the one people forget. An agent works in a loop: Claude asks for a tool (a function your code offers it, such as "read this file"), your code runs it and sends back the result, and Claude carries on. A batch request has no open line back to your code while it runs. If Claude's reply asks for one of your tools, the request ends right there, and its result is the tool request itself (the reply's stop_reason field, which says why Claude stopped, reads "tool_use"). Carrying on would take a brand-new request with the tool's answer added.
So a batch request must carry everything Claude needs up front. That is why the test-gap report can be batched: each request already contains the module's code and its existing tests. A Claude Code review that reads files and runs commands as it goes is the opposite case. Each of those steps is a tool call that Claude Code executes on your side, so the review could not live inside one batch request.
4.5.3 Match the API to who is waiting
One question settles every batching proposal, and it isn't about savings. It's this: who is blocked until the answer arrives? The latency requirement decides first; the discount only applies to work that passes that test.
The pre-merge check fails it. A developer has finished a pull request and can't merge until the review is back. If the review usually arrives within the hour but may take a whole day, the team has swapped a one-minute wait for an unpredictable one, and merges pile up behind it.
Two patches look tempting, and neither works. Polling (your code asking "is it done yet?" every few seconds) notices sooner that a batch has finished, but it doesn't make the batch finish sooner. A timeout that falls back to a real-time call when the batch is slow adds a second code path to build and test. Its worst case is a slower version of the synchronous call you started with.
The test-gap report passes easily. It's latency-tolerant: it runs while nobody is looking, and if it's late one morning, people read it when it lands. Nobody's work is blocked. The same holds for overnight reports of any kind, weekly audits and nightly test generation: work that runs on a schedule and is read later. These are non-blocking workloads. Batch them and take the 50%.
Two workflows, two APIs
Pre-merge check blocking
Test-gap report overnight
So the right answer to her proposal is neither yes nor no, but a split: move the report, keep the check.
4.5.4 Plan submissions around the 24-hour ceiling
Now a harder version. The report works so well that the QA lead makes the engineering teams a promise, an SLA: every merged pull request gets its test-gap analysis within 30 hours of the merge. Merges arrive all day, so the pipeline collects them and submits a batch on a schedule. How often must it submit?
The trap is to plan with the typical time: "batches usually take an hour, so one a day is plenty." A promise has to survive the worst case, and the worst case has two parts. First, a pull request merged just after a submission waits for the next one: up to one full interval. Then that batch may take the whole 24 hours. The interval plus 24 hours must fit inside the 30.
| Submission interval | Worst case for one pull request | Verdict |
|---|---|---|
| Once a day | 24 h waiting + 24 h processing = 48 h | SLA broken |
| Every 6 hours | 6 h + 24 h = 30 h | Kept only on paper: no time left to read and publish the results |
| Every 4 hours | 4 h + 24 h = 28 h | SLA kept, with 2 hours of margin |
Submitting in 4-hour windows is the answer. The unluckiest pull request is analysed within 28 hours, and the 2 hours that remain cover downloading the results and writing them into the report. The general rule: the submission interval can be at most the SLA, minus the 24-hour window, minus whatever time your own steps need. The typical one-hour batch never enters the calculation.
4.5.5 custom_id and the batch life cycle
Here's a worry the exam likes to plant: "batch results come back out of order, so how would we know which test gaps belong to which module?" The worry is half right. The docs say plainly that results can arrive in any order. The answer isn't to avoid batching; it's the custom_id.
Every request in a batch carries a custom_id that you choose: 1 to 64 characters of letters, digits, hyphens and underscores, unique within the batch. It's the name tag on each job in the print shop's box. Anthropic returns it untouched on the matching result, so the order results arrive in stops mattering. Make it meaningful: tests-payments or tests-auth-service tells you at a glance which module a result belongs to. Here is the report's batch, start to finish.
- BUILD. One request per module, each with its
custom_idand the ordinary Messages API settings: the model, the instructions, and the module's code and tests. The instructions ask for the gaps as JSON, a structured format another program can read. - SUBMIT. Send the list. You get a batch ID back at once, with
processing_statusset toin_progress. - POLL. Every minute or so, your code asks for the batch by its ID until
processing_statusreadsended. Therequest_countsfield tallies how many requests are still processing, succeeded, errored, were canceled or expired. - READ. Fetch the results: one line per request, each carrying its
custom_idand a result type. Your code joins each result to its module bycustom_id. - RESUBMIT. Collect the
custom_ids that didn't succeed and send only those in a new batch.
The life cycle of a batch
custom_idprocessing_status is endedcustom_id4.5.6 When some requests fail
The first full run of the test-gap report covers 400 modules and ends with 392 succeeded, 5 errored and 3 expired. The tempting move is to fix whatever needs fixing and send all 400 again. Resist it. You would pay a second time for 392 answers you already have, and wait up to another 24 hours for them. Each result type tells you exactly what to do with that one custom_id.
| Result type | What happened | What to do with that custom_id |
|---|---|---|
succeeded |
Claude replied; the message is included | Use it (after your usual validation) |
errored |
No reply: the request was invalid, or a server error occurred | Invalid: fix it first. Server error: resubmit it unchanged |
canceled |
You canceled the batch before this request ran | Resubmit it if you still need it |
expired |
The 24-hour window closed before this request ran | Resubmit it unchanged |
Only succeeded requests are billed; errored, canceled and expired ones cost nothing. Resending the whole batch pays twice for every success, while resending only the failures costs just what they use.
Look closer at the 5 errored modules. Their error is an invalid request. Each is a huge legacy module whose code alone is longer than the model's context window (the most text the model can take in at once). The API rejects any request whose input exceeds it. Resend them unchanged and they fail again, identically. The appropriate modification is chunking: split each module's files into parts that fit, give every part its own ID (tests-billing-legacy-part1, -part2 and so on), and merge the parts' results when they come back. The 3 expired requests just ran out of time, so they go back as they were.
The cheapest failure, though, is the one you never submit. Before any full run, the team tries the prompt on a sample of 20 modules, hard cases included, and reads every output. They fix what's wrong: gaps listed without naming the untested behaviour, tests drafted for generated code nobody maintains. A flaw that shows up in 10% of the sample shows up in about 40 modules at full volume. Each of those wrong answers is billed (a bad answer still counts as succeeded) and then has to be sent again.
Refining on a sample first maximises the first-pass success rate: the share of requests that come back right the first time. The docs add a smaller reason. A batch checks each request's settings only while it processes them, and reports those errors only when the whole batch has ended. So send one request through the synchronous API first, to confirm it is well formed, before you send thousands.
4.5.7 The exam traps
Every trap below is a batching decision made on price or habit instead of on the facts of the workload.
- ✗ Moving a blocking pre-merge check to the batch API "with polling" or "with a timeout fallback". ✓ Keep it synchronous. With no latency SLA, "usually within an hour" can't carry a workflow people wait on; polling speeds nothing up, and a fallback adds complexity to rescue a choice you didn't need to make.
- ✗ Keeping latency-tolerant work synchronous because batch results arrive out of order. ✓ Batch it and match results by
custom_id. Ordering is a solved problem, not a reason to pay double. - ✗ Setting the submission schedule from typical batch times. ✓ Plan with the worst case: the interval plus 24 hours must fit the SLA. Daily submissions break a 30-hour SLA; every 4 hours keeps it.
- ✗ Resubmitting the whole batch after a few failures. ✓ Resubmit only the failed
custom_ids, modifying the ones that would fail again, such as chunking a document that exceeded the context limit. - ✗ Batching an agent that calls your tools as it works. ✓ Your code can't execute a tool and return its result in the middle of a batch request. Put everything the model needs into the request, or keep the agent synchronous.
- ✗ Sending the full volume with an untested prompt. ✓ Refine the prompt on a sample set first; every flaw found at full volume is paid for, and waited for, again.
4.5.8 Put it together: run a batch and break it
You now have every piece. You know what the discount costs, which work can wait and how an SLA becomes a submission schedule. You know how custom_id holds a batch together and how to handle failures without paying twice. A small build makes it stick, mostly because you'll feel the wait.
The next task statement, multi-instance and multi-pass review (4.6), splits a large review into focused per-file passes plus integration passes, and adds an independent Claude instance that reviews the generator's work. Every pass that runs as a self-contained request raises this lesson's question again. If a developer is waiting on it, it runs synchronously; if its findings are read in the morning, it can go into a batch at half the price.
Key takeaways
- ✓ The Message Batches API costs 50% less than the synchronous API, but processing can take up to 24 hours and there is no guaranteed latency SLA.
- ✓ Your code can't run its tools mid-request and return the results inside a batch request, so multi-turn tool calling stays on the synchronous API.
- ✓ Choose by who is waiting: blocking work such as a pre-merge check stays synchronous; overnight reports, weekly audits and nightly test generation go to a batch.
- ✓ For an SLA, the worst case is the submission interval plus 24 hours, so submitting every 4 hours keeps a 30-hour SLA.
- ✓ Results can arrive in any order; the unique
custom_idon every request is how you match them. - ✓ Resubmit only the failed
custom_ids, modified where needed, such as chunking a document that exceeded the context limit. - ✓ Refine the prompt on a sample set before the full batch, so the first pass succeeds and resubmissions stay rare.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
72 CCAR-F questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.