Home › Study guides › CCAR-P › Domain 3 › Lesson 3.3
CCAR-P · Domain 3 · 19% of the exam · Lesson 3.3 · 22 min read
Accuracy versus latency: measure both, then justify the configuration
How to measure time to first token, p95 latency and accuracy on one eval set, which levers trade one for the other, and how to justify what you ship.
Written against objective 3.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.3.1 When the right answer arrives too late
A winter storm closes Corvina Air's hub for six hours. Four hundred flights are cancelled, and the disruption desk's phone queue passes two thousand callers. On every call the agent finds seats and then works out what the ticket allows: whether the storm waiver covers a free change, whether a fare difference is owed, whether a partner airline's flight counts. The passenger stays on the line throughout, listening to the silence.
Saoirse, a solution architect, was brought in to build an assistant that proposes rebooking options to the agent during that silence. Her pilot ran a top-tier model told to reason at length over forty chunks, short passages of tariff and waiver text, on every call. A second call checked each fare rule before anything reached the screen. It was right on 98 calls in 100. On a busy morning it also took twelve seconds to show a single word. Leandro, who runs the desk, watched agents abandon it by lunchtime and go back to quoting waivers from memory, which was far less reliable.
The obvious swap failed in a different way. The smallest, fastest model with a handful of chunks answered in about a second and got the fare rule wrong on more than one call in eight. Wendell, who leads revenue integrity, counts each of those as a refund paid that should not have been, or a fare difference never collected. Neither pilot failed for want of a better model. Each one fixed a single number and ignored the other.
Accuracy and latency are properties of a whole configuration. That covers the model tier, how much it thinks and retrieves, how many calls you chain, what runs in parallel, what is cached and how the answer is delivered. Every setting moves both numbers, so the architect measures both on the same cases and chooses between them deliberately.
3.3.2 Measure what the user actually waits for
"How fast is it?" derails most latency reviews, because a reply that streams has two moments that matter. Time to first token (TTFT) is the time from sending the request until the first token of the reply arrives. End-to-end latency runs until the last token, when the answer is complete. For a whole pipeline, both clocks start when the user acts, so the lookups and retrieval before the model call count too.
The gap between those clocks is the gap between perceived and actual latency. With streaming, the API sends the reply as it is generated, so the agent can read out the first option while the rest is still being written. Think of a kitchen that serves each course as soon as it is plated: no faster cooking, but the table starts eating sooner. One catch: Claude can think before it answers, in a thinking block ahead of the reply that streams back empty by default on Claude Opus 5.5 and Sonnet 5.5. So Corvina's headline number is time to the first VISIBLE text at the agent's screen, not to the first streamed event.
Where the seconds go in one rebooking request
Report each clock as percentiles, not an average. The p50 is the median: half the requests are faster. The p95 is the time that one request in twenty exceeds. An average hides exactly the calls that hurt. A desk taking two thousand calls an hour with a p95 of eight seconds leaves a hundred callers an hour in eight seconds or more of dead air. Measure under peak load, too: a quiet Tuesday and a storm morning are different systems.
Accuracy needs the same discipline. Build an eval set, real cases with known correct answers, run every configuration over it, and score the error that costs money on its own rather than inside a general quality score. Saoirse replayed 240 disruption calls from last winter, each with the correct fare treatment signed off by Wendell's team. The requirement she agreed fits in three lines; look at where each number is measured and on which cases.
Latency: time to first visible text at the agent's screen, p95 at or below 2.0 s during a peak disruption hour, with p50 reported alongside. All options on screen: p95 at or below 6 s.
Accuracy: share of the 240 replayed disruption calls in which every proposed option applies the fare and waiver rules correctly, reported for the whole set and for partner-airline itineraries separately.
Method: latency and accuracy come from the same run over the same cases, for every configuration considered.
3.3.3 Take the latency that costs no accuracy first
It is tempting to open a latency problem by shrinking the model. Resist it until you have collected the seconds that cost nothing. Anthropic's latency guide sets the order: first get a prompt working well without constraints, then reduce latency, because cutting early hides what top performance looks like. So Saoirse's slow pilot was the right start: it showed the accuracy ceiling, about 98%.
Four changes cut her latency without changing the evidence the model sees or how hard it reasons. Streaming puts the first option on screen while the rest is generated. Prompt caching lets the API reuse its processed form of a prompt's opening, its prefix, on later requests that start the same way, which cuts processing time and cost. The tool definitions, instructions and waiver policy go first and are cached; the passenger's booking goes last.
Parallel lookups fetch the fare rules and the seat map at the same time, so the wait is the slower of the two, not their sum. When the model asks for several independent tools in one reply, your code can run those together as well. Finally, a compact output format (three options, one line each, the fare treatment stated plainly) removes tokens the agent never needed. On the same 240 calls, they cut the pilot's p95 time to first text from 11.8 to 8.7 seconds, and accuracy did not move.
The same answer, delivered sooner
Pilot pipeline
After the free cuts
Caching has conditions an architect must check, because a miss costs time without raising an error. A prefix shorter than the model's minimum is not cached at all. An entry lives five minutes by default, refreshed on each read, so a quiet night can leave the first calls of a storm uncached. A pre-warm request with max_tokens: 0 loads the prefix before the surge. The thinking configuration and the effort level (how hard the model is asked to reason) are rendered into the prompt, so changing either starts the cache over. Pick one configuration per route and hold it.
3.3.4 The levers that trade accuracy for time
The remaining 8.7 seconds were high-effort thinking over forty chunks, plus a whole second call in front of the first word. Every lever from here buys accuracy with time, or time with accuracy, and the table is the architect's menu.
| Lever | When it earns its latency | What it costs |
|---|---|---|
| Model tier | Your eval shows a smaller tier failing on reasoning the task needs | Comparative latency runs from Haiku (fastest) through Sonnet and Opus to Fable (slower); real latency also depends on prompt, output and effort |
| Thinking and effort | Errors come from rules that interact: date windows, exceptions, combinations | Thinking happens before the first text, so more of it delays the first word; lower effort is the first lever on a thinking model |
| Retrieval depth | The rule that decides the answer often sits outside the top few chunks | Every extra chunk is more input to process and more noise to read past |
| Reranking | Many candidates, few relevant: a second model rescores a wide set and keeps the best handful | A small extra step at runtime, offset in part by sending the model fewer chunks |
| Verification pass | One checkable error is expensive, such as a wrong fare rule | A full extra call; placed before the first text, it adds its whole duration |
| Chained steps | The task splits into fixed, easier subtasks the model gets right more often | Every step adds a round trip and its own time to first token |
On current models, effort (output_config.effort on the request) is the main dial for thinking, and it shapes every output token. Claude Sonnet 5.5 defaults to high, and Anthropic's guidance for chat and other latency-sensitive work is to start at medium or low. Claude Haiku 4.5 has no effort setting and thinks only when given a manual thinking budget. Effort is guidance, not a cap: at low effort the model still thinks on a hard problem, only less, which suits a desk where most calls are simple.
Before pulling any lever, Saoirse read the fast pilot's failures. About two thirds were waiver clauses that never reached the model, because the relevant text sat outside the top five chunks. The rest were misread date windows: a waiver covering travel "within seven days of the original flight" applied as seven days from today. A bigger model fixes none of the first kind; it cannot reason about a clause it never sees. So she spent latency in two places: wider retrieval with reranking for the missing clauses, and a thinking model at low effort for the dates. Anthropic's retrieval experiments name that trade plainly: rerank more chunks for accuracy, fewer for speed.
3.3.5 Measure the candidates, then choose
Saoirse built four candidates, all with the free cuts, and ran each over the same 240 calls at peak-hour concurrency. Look at the loop over text_stream, which times the first visible text rather than the first event, and at cache_read_input_tokens, which proves the static prefix really came from the cache.
import time
import anthropic
client = anthropic.Anthropic()
CONFIGS = { # retrieval, reranking and the fare check live in the pipeline around this call
"A": {"model": "claude-opus-5-5", "output_config": {"effort": "high"}},
"B": {"model": "claude-haiku-4-5-20251001"}, # no effort parameter; thinking off by default
"C": {"model": "claude-sonnet-5-5", "output_config": {"effort": "low"}},
}
def timed_run(name: str, system: list, case: str) -> dict:
start, first_text = time.perf_counter(), None
with client.messages.stream(max_tokens=16000, system=system,
messages=[{"role": "user", "content": case}], **CONFIGS[name]) as stream:
for _ in stream.text_stream: # text deltas only, never an empty thinking block
first_text = first_text or time.perf_counter() - start
final = stream.get_final_message()
return {"first_text_s": first_text, "complete_s": time.perf_counter() - start,
"cache_read": final.usage.cache_read_input_tokens, # 0 means the prefix missed
"answer": next(b.text for b in final.content if b.type == "text")}
The harness times the model call alone; the table's figures were taken at the agent's screen, so they include lookups and reranking. Candidate D reuses C's call and adds a narrow fare check that starts as soon as the options are complete, while the agent is still reading them out. The confirm button stays locked until it passes, and a flagged option goes back for a corrected fare.
| Configuration | First text, p50 / p95 | All options, p95 | Fare rules right |
|---|---|---|---|
| A. Opus 5.5, high effort, 40 chunks, check before display | 5.2 s / 8.7 s | 13.9 s | 98.3% |
| B. Haiku 4.5, no thinking, top 5 chunks, no check | 0.6 s / 1.1 s | 3.2 s | 86.7% |
| C. Sonnet 5.5, low effort, 50 chunks reranked to 8 | 1.0 s / 1.7 s | 5.1 s | 94.6% |
| D. C plus a fare check that gates confirmation | 1.0 s / 1.7 s | 5.1 s | 97.9% |
Now plot them: p95 first text along the bottom, accuracy up the side, a vertical line at two seconds. The requirement that cannot bend is a filter, not a score, so everything right of the line is out, and among the rest you take the highest point. It is how you hire for a role that needs a licence: the licence filters the candidates, then you pick the strongest who holds one.
Filter by the requirement, then maximise accuracy
Misses 2.0 s at p95
Meets 2.0 s at p95
Sometimes the fixed requirement is accuracy. A dosing check or a sanctions screen has an accuracy floor: remove everything below it, then take the fastest survivor. If nothing passes, that is a finding, not a configuration: take the numbers to the stakeholders and change the design (moving work off the critical path, as D does) or change the requirement openly.
Finally, look inside the totals. D scores 97.9% overall but 91.0% on partner-airline itineraries, about 6% of calls, and p95 latency deserves the same split by call type. A weak segment whose users can wait may get its own route: a slower, more careful configuration for those cases only, measured the same way. Corvina's passengers cannot wait, so partner itineraries show the agent a "verify fare" flag instead.
3.3.6 Justify the choice in a decision record
A configuration nobody can explain gets reversed at the first complaint. Wendell will ask why the desk does not run the most accurate option; Leandro will ask why it is not the fastest. A short decision record answers both in advance, in five parts: the requirement, the options measured, the choice, the residual risk and what would reopen it. Here is Saoirse's.
Decision: configuration of the rebooking assistant on the disruption desk. Owner: Saoirse. Date: 30 September 2026.
Requirement: first visible text within 2.0 s at p95 at the agent's screen during a peak hour (Leandro). Fare and waiver rules right on as many calls as possible, since each error is a refund paid or a fare difference lost (Wendell).
Options measured on the same 240 calls under the same load: A, Opus 5.5 at high effort with a check before display: 98.3%, 8.7 s. B, Haiku 4.5 with 5 chunks: 86.7%, 1.1 s. C, Sonnet 5.5 at low effort with reranked retrieval: 94.6%, 1.7 s. D, C plus a fare check that gates confirmation: 97.9%, 1.7 s.
Chosen: D, the most accurate configuration that meets the latency requirement. A is 0.4 points more accurate but more than four times the limit at p95, and agents abandoned it in the pilot.
Residual risk: about 2 calls in 100 still carry a fare error the check misses. Partner-airline itineraries score 91.0% and show the agent a "verify fare" flag. A waiver published mid-storm is invisible until the index refreshes.
Revisit when: p95 first text exceeds 2.0 s for a week, weekly sampled accuracy falls below 97%, a new model is released, or the waiver format changes.
Notice what makes it defensible. Every option is measured on identical evidence, so nobody can argue from a benchmark or a hunch. The 0.4 points given up for speed are stated, not hidden. The residual risk is named, with a mitigation for the weakest segment, and the triggers turn "revisit it someday" into conditions a dashboard can watch.
The record is also where a weak justification shows up. "Sonnet is faster" is not a reason. "C and D meet the 2.0-second p95 at the desk, and D is 3.3 points more accurate than C" is. If a sentence in the record could have been written without running the eval, it is opinion, and the stakeholder who disagrees has an equal claim.
3.3.7 The exam traps
Every trap here optimises one number without measuring the other, or fixes the wrong step.
- ✗ Reporting one average latency figure. ✓ Report time to first visible text and time to a complete answer at p50 and p95, under peak load, where the user waits. The average hides the calls that hurt.
- ✗ Turning on streaming to fix a wait that happens before the first token. ✓ Streaming changes perceived latency only after the first token. Seconds spent in sequential lookups, an uncached prefix or heavy thinking must be fixed at their source.
- ✗ Downsizing the model as the first move against latency. ✓ Take the cuts that cost no accuracy first (cache the static prefix, run independent lookups in parallel, stream, trim the output), then trade with eval evidence.
- ✗ Buying accuracy with the biggest model and maximum effort on every request. ✓ Read the failures and spend latency where they are: retrieval misses need retrieval and reranking, and a verification pass belongs off the critical path where the product allows.
- ✗ Hitting the number by cutting blindly: fewer chunks, a lower
max_tokens. ✓ Rerank to fewer, better chunks and ask for a compact format.max_tokensis a hard stop that can end an answer mid-sentence. - ✗ Choosing on a public benchmark or a hunch and leaving the choice unwritten. ✓ Measure every candidate on your own eval set, and record the requirement, options, choice, residual risk and triggers.
3.3.8 Put it together: measure two configurations and defend one
You now have the whole method. Two clocks and a fixed eval set, the free latency first, the levers that trade, a choice filtered by the fixed requirement, and a record that defends it. Own it with a small run of your own, broken on purpose and written up from your own numbers.
The rest of the domain keeps these two numbers honest. Observability at scale (3.4) measures the p95 and the weekly accuracy sample in production, per step and per configuration. RAG pipeline design (3.5) and retrieval strategies (3.6) decide what a chunk is and how candidates are found, which caps what reranking can achieve. Later, A/B testing (4.3) checks on live traffic what the offline eval predicted, and cost-performance optimisation (4.5) adds the third axis.
Key takeaways
- ✓ Accuracy and latency are both outputs of one configuration, so measure them together, on the same eval set and in the same run.
- ✓ Latency is two clocks, time to first visible text and time to a complete answer, reported at p50 and p95 under peak load where the user waits.
- ✓ Streaming, prompt caching, parallel independent lookups and a compact output cut latency without costing accuracy; take them first, after finding the accuracy ceiling.
- ✓ Model tier, effort, retrieval depth, reranking, verification passes and chained steps trade time for accuracy, so spend latency where the eval's failures are.
- ✓ Choose by filtering on the requirement that cannot bend and maximising the other number among the survivors, checking segments as well as totals.
- ✓ Justify the choice in a decision record: requirement, options measured, choice and reason, residual risk, and the triggers that reopen it.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAR-P questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Integration alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.