Home › Study guides › CCAR-P › Domain 3 › Lesson 3.4
CCAR-P · Domain 3 · 19% of the exam · Lesson 3.4 · 23 min read
Observability at scale: tracing, sampling and scoring LLM traffic
Why an LLM system fails with HTTP 200, what to trace in every request, and how to sample, redact and score millions of conversations at a sane cost.
Written against objective 3.4 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.4.1 Why the complaints arrive before the alerts
Norrvik Telecom's care platform handles about two million conversations a month. Customers chat with Claude in the app, and in the contact centre an agent-assist tool drafts replies for human care agents. Behind both sits one pipeline. A small, fast model sorts each message by intent, a retrieval step finds tariff and policy articles, and a mid-tier model writes the answer, calling tools such as get_account and check_network_status. The most capable model takes the hardest billing disputes.
In March, Norrvik relaunched its roaming bundles. Six weeks later the complaints began: customers had been told a bundle covered countries it did not, and their bills said otherwise. Oona, the architect brought in to find out why, opened the dashboards first. They were green throughout: 99.9% of requests succeeded, p95 latency (the time within which 95% of requests finish) never moved, and spend was on forecast.
Then she looked for the conversations and found two kinds of record, neither usable. The service logs held a status code, a duration and a token count per request, and nothing about which articles were retrieved or which prompt answered. A debugging index an engineer had switched on held full transcripts, phone numbers and addresses included, searchable by two hundred staff. Ilse, Norrvik's data protection officer, wanted it gone that week.
The problem is not a missing tool. An LLM system fails in a way request monitoring was never built to see: a fast, well-formed, successful response that happens to be wrong. Observability for such a system means recording what happened INSIDE each request (every step, version and token) and judging the answers themselves on a sample. It must do both without hoarding personal data or inflating the telemetry bill. Oona's job is to decide what is traced, sampled, redacted and scored.
3.4.2 What makes an LLM system hard to observe
Why doesn't the monitoring that already runs Norrvik's billing and network services do the job? Those services fail loudly, and classic monitoring assumes that a bad outcome shows up as an error code, a timeout or a slow response.
An LLM system breaks that assumption in six ways, and each one creates a design requirement.
| Property | What request monitoring shows | What the design needs |
|---|---|---|
| Quality failures return HTTP 200 | A fast, successful response | A quality signal on answers: scores and feedback |
| Outputs vary from run to run | A replay that gives a different answer | A record of what was sent and returned at the time |
| One request spans many steps | One duration for the whole request | A span per step, joined by one trace id |
| Cost varies per request | A monthly bill per model | Tokens on every model call, by route and prompt version |
| Payloads contain personal data | A log store that copies every conversation | Content redacted before storage, kept only where needed |
| Volume is high | Storage that costs more than it tells you | Metrics on all traffic, traces kept by rule |
You will usually meet the six as symptoms, not names: green dashboards, a replay that works fine, a bill nobody can explain. The second defeats the oldest debugging habit, sending the same request again to watch it fail. Replaying a roaming question today proves nothing about what a Norrvik customer read in March, when a different prompt was live. Anthropic hit the same wall with its multi-agent Research feature. Its agents were non-deterministic between runs even with identical prompts, so reports of agents "not finding obvious information" stayed unexplained until full production tracing showed what each run had done.
3.4.3 What to record: one trace per turn, one span per step
The natural unit is the trace, the record of one request's path through every service. It is made of spans, one per step, each with a duration and attributes. At Norrvik a trace covers one conversational turn. Every span in it shares one trace id, and a conversation id links the turns.
One conversational turn as a trace
Root span one turn
Model and retrieval spans
Tool and outcome spans
On every model call, record the model id, prompt version, input and output tokens, cache reads and writes, stop_reason (why generation ended), latency, status code and the API's request-id. On retrieval, record the index version and the ids and scores of the chunks returned, not their text. On every tool call, record the name, outcome and duration.
In the model-call span below, look at the attributes read from response.usage, and at the request_id line, which gives Anthropic support a handle on this exact call.
from opentelemetry import trace
import anthropic
tracer = trace.get_tracer("norrvik.care")
client = anthropic.Anthropic()
def answer_turn(messages, route, prompt_version):
# Child of the turn's root span; an exception leaving the block marks the span as an error
with tracer.start_as_current_span("llm.answer") as span:
span.set_attribute("app.route", route) # bounded: "roaming", "billing"
span.set_attribute("app.prompt_version", prompt_version) # "answer-v14"
response = client.messages.create(model=ANSWER_MODEL, max_tokens=1024,
system=PROMPTS[prompt_version], messages=messages)
usage = response.usage
span.set_attribute("llm.model", response.model)
span.set_attribute("llm.input_tokens", usage.input_tokens) # after the last cache breakpoint only
span.set_attribute("llm.cache_read_tokens", usage.cache_read_input_tokens or 0)
span.set_attribute("llm.cache_write_tokens", usage.cache_creation_input_tokens or 0)
span.set_attribute("llm.output_tokens", usage.output_tokens)
span.set_attribute("llm.stop_reason", response.stop_reason) # watch max_tokens and refusal
span.set_attribute("anthropic.request_id", response._request_id) # quote it to support
return response
Two details catch teams out. First, with prompt caching on, input_tokens counts only the tokens after the last cache breakpoint. A span that records it alone makes a cached 6,000-token policy prefix vanish from your cost figures, so record the cache fields too. Second, the official SDKs retry rate-limit (429) and server errors twice by default, so a span around the call sees only the final outcome. If rate-limit pressure matters, lower the SDK's max_retries and retry in your own code with a span per attempt.
The trace id must cross every boundary, the services behind the tools included. OpenTelemetry (OTel), the vendor-neutral standard for traces, metrics and logs, carries it between services in the W3C traceparent header, and inside the message when one agent hands work to another. Making OTel the common format means each team instruments once, and the backend stays a separate, replaceable choice. Claude Code and the Agent SDK already export through it, with spans (in beta) shaped like the ones above. There, a subagent's spans nest under the tool call that launched it, so a delegation chain reads as one trace.
One warning from the Agent SDK docs applies to any design: its identity attributes name your service's credential, not the customer the agent acted for. So Oona puts a pseudonymous customer id on every root span, which lets a complaint that arrives six weeks late be joined, through a restricted lookup, to the turns behind it.
3.4.4 Sampling, redaction and the cost of seeing
Two million conversations a month, at about six turns each and eight spans a turn, is close to a hundred million spans a month. Add the text of prompts, passages and answers, and you have built a second copy of every customer conversation. Tracing backends usually bill by volume, and the privacy team must defend every copy.
Three moves keep it in proportion. The first is metrics on everything: counters and latency histograms computed from every request before any sampling, labelled only with bounded values such as model, route, prompt version, tool and stop_reason. A customer id as a metric label creates one time series per customer, the cardinality problem, so ids belong on traces.
The second is tail sampling: deciding whether to keep a trace after it completes, when you know how it ended. Norrvik keeps every trace with an error, a 429 (rate limit) or 529 (overloaded) response, a refusal or max_tokens stop, an escalation, negative feedback or a care-agent edit. It also keeps every trace whose cost is above the 99th percentile, plus 2% of the rest at random. That 2% is not waste; it is the baseline that failures are compared against, and the pool that quality scoring draws from.
The third is redaction before storage. Structure (durations, model names, token counts, ids) goes on every span. That is also the Agent SDK's default: prompt text and tool content are opt-in, and its docs advise leaving them off unless your pipeline is approved to store that data. Norrvik needs content to diagnose and grade answers, so its spans carry it to the collector, the service that receives telemetry before storage. The collector masks phone, account and card numbers and addresses before anything else reads them, and only the content of kept traces reaches a separate store with named readers and 30-day retention. Masking inside each service would be stricter; Norrvik chose the collector so that one set of rules covers every team.
Norrvik's telemetry pipeline
Sampling is where designs most often go wrong.
| Option | When it wins | What it costs |
|---|---|---|
| Head sampling (keep a random share, decided when the request starts) | High volume where failures are common enough to catch at random | Keeps the same small share of failures as of successes, so rare failures vanish |
| Tail sampling (decide after the trace ends) | High volume where failures are rare and matter | A collector that buffers whole traces before it decides |
| Metrics only | Capacity planning and finance reporting | No way to explain any single answer |
The requirement that decides is how rare and how costly the failures are. At Norrvik a wrong roaming answer might be one turn in a thousand. Sampling 1% at the start keeps about one in a hundred of those, so the trace behind a complaint is almost never there. Think of a shop's security camera: recording a random 1% of the day almost never catches the theft, while keeping a rolling buffer and saving the footage around every incident always does. Tail sampling is the rolling buffer: it keeps every failure.
3.4.5 Watching quality, not only uptime
The pipeline now keeps the right traces, but none of them says whether an answer was right. That takes three kinds of quality signal: automated scores, feedback from people, and the mix of what people ask.
Automated scoring of sampled traffic comes first. Cheap code checks run on every answer: does every price it quotes appear in the retrieved passages, did it call get_account before quoting a balance? Judgement calls go to an LLM grader, a separate model call that scores an answer against a rubric. Anthropic's evaluation guidance gives the recipe: a clear rubric, a specific output such as a 1-to-5 score, reasoning before the score, and a reliability test before scaling up, for example against answers people have scored. The grader runs off the request path, so it can use the Message Batches API at half the price.
The core of Norrvik's grader prompt ties every level to evidence in the retrieved articles, not to how confident the answer sounds.
Roaming answers: read the customer's question, the retrieved articles and the answer. Reason step by step inside <thinking> tags, then give <score>1-5</score>.
5: Every country, price and allowance in the answer appears in the retrieved articles, and the answer names the bundle it applies to.
3: The answer is correct but leaves out a condition stated in the articles, such as a daily cap or an end date.
1: The answer states a country, price or allowance that the articles do not support, or contradicts them.
If the articles do not answer the question, an answer that says so and offers a care agent scores 5.
Feedback signals from people come next. A thumbs-down in the app is explicit but rare. Implicit signals are richer: the customer contacts care again within seven days on the same topic, or asks for a person. In agent-assist, the best signal is what the care agent does with a draft: send it, edit it or discard it. These arrive within days, not six weeks.
Drift in the input mix is the third. A quality score is an average over whatever customers asked this week, so it moves when the questions move. A bundle launch or a network outage shifts the mix of intents overnight. Norrvik tracks each intent's share and the rate of retrievals with no good match against a four-week baseline, and reports every score per intent and prompt version.
Where quality signals come from
Replay March through this design and the roaming failure surfaces in days. Roaming questions tripled after the launch, so the overall average blended a change in what people asked with a change in how well they were answered. Split by intent, roaming scores fell while every other intent held steady. Split by prompt version, the low scores all came from answer-v14, released with the new bundles, which had dropped the instruction to quote allowances only from retrieved articles. Care agents' edit rate on roaming drafts had doubled.
3.4.6 Choosing a strategy for the scale and the constraint
No single design fits every system. Four requirements decide most cases: how much traffic there is, how sensitive the payloads are, how many services and agents one request crosses, and how quickly a quality problem must be found.
| Situation | Strategy that wins | What it costs |
|---|---|---|
| Pilot, a few thousand conversations, no personal data | Keep every trace with content; review answers by hand | Cost and exposure grow with traffic; does not survive scale-up |
| High volume with personal data, like Norrvik | Metrics on all traffic, tail sampling, redacted content on kept traces, sampled scoring | A collector, redaction rules to maintain, a grader to validate |
| Content may not be stored at all (a contract or policy forbids it) | Structural spans only; score answers in memory, keep the scores, discard the text | Slower diagnosis of the hard cases |
| Many services and agents per request | Shared trace context in every service, before anything else | Instrumentation work in every team, tool owners included |
| Finance needs spend by team and model | The Usage and Cost Admin API | Aggregates by API key, workspace and model, never per request |
The content-free row is a real option: Anthropic's Research team monitors agent decision patterns and interaction structures without monitoring the contents of individual conversations. The last row is right for finance and a trap for diagnosis: those reports reconcile the bill but cannot say which flow or prompt caused a rise.
An architect justifies the choice in writing: the requirement each part serves, and the options rejected. Here is Norrvik's record.
Decision: Norrvik care telemetry, version 1.
Context: About 2 million conversations a month; payloads hold phone numbers, account numbers and addresses; wrong answers were discovered six weeks late.
Requirements: Detect a quality regression in any intent within three days. No raw personal data outside the restricted store.
Tracing: One trace per turn; spans for classification, retrieval, each model call and each tool call; trace context propagated to every service; pseudonymous customer id on the root span.
Keeping: Metrics on all traffic. Tail sampling keeps every failure, flag and cost outlier plus 2% of the rest. Content on kept traces only, redacted in the collector, 30-day retention.
Quality: Code checks on every answer; batch LLM grader on kept traces, validated against 300 human-scored answers; care-agent edit rate and 7-day re-contact per intent; intent mix against a four-week baseline.
Rejected: Full transcript logging (privacy, cost). Metrics only (cannot explain one answer). 1% head sampling (loses most rare failures).
Notice that it picks no vendor, draws no dashboard and changes no model: every line maps to a requirement someone can check.
3.4.7 The exam traps
Every trap here either watches the wrong thing or watches the right thing at the wrong cost.
- ✗ Judging the system's health by success rate and latency. ✓ Add quality signals: scores on sampled answers, feedback and the input mix per intent. A wrong answer returns 200 and arrives on time.
- ✗ Logging every prompt and answer in full, "so we can debug anything". ✓ Keep structure on every span and content only on kept traces, redacted before storage. Full logs are a second copy of every conversation, with its own cost and privacy duty.
- ✗ Cutting telemetry cost with 1% head sampling. ✓ Count metrics on all traffic and tail-sample, keeping every failure plus a small random baseline. Head sampling drops rare failures at the same rate as successes.
- ✗ Explaining per-request cost from the organisation's usage and cost reports. ✓ Record tokens, cache reads and writes included, on every model-call span with route and prompt version. The reports aggregate by key, workspace and model.
- ✗ Letting each service keep its own logs under its own ids. ✓ Propagate one trace context across services, agents and tools. Without it, nobody can connect a slow tool to the answer it delayed.
- ✗ Moving to a bigger model when an average quality score drops. ✓ Split the score by intent and prompt version and check the input mix first. The cause is usually a new prompt or a new kind of question, and the traces show which.
3.4.8 Put it together: design the telemetry for one flow
You now have every piece, from the six properties to the requirements that pick a strategy. The quickest way to own them is to instrument one small flow, break the knowledge behind it, and see which signals notice.
The rest of the guide builds on this telemetry. Evaluation metrics (4.1) define, before release, the rubric your live grader applies after it. Monitoring with logging and observability tools (4.6) turns these traces and metrics into dashboards, alerts and incident work. Compliance (5.4) sets the retention and access rules your redacted store must meet.
Key takeaways
- ✓ LLM systems fail quietly: wrong answers return HTTP 200, outputs vary between runs, and one answer hides many steps, so uptime monitoring misses the failures that matter.
- ✓ Record one trace per turn with a span per step, and put the model id, prompt version, tokens including cache reads, stop reason, latency, status and request id on every model call.
- ✓ Propagate one trace context across every service, agent and tool, and add a pseudonymous customer id so a late complaint can be joined to its trace.
- ✓ At scale, count metrics on all traffic with bounded labels, tail-sample traces to keep every failure plus a random baseline, and redact content before it is stored.
- ✓ Measure quality directly with code checks, a validated LLM grader on sampled traffic, feedback from people and the input mix, reported per intent and prompt version.
- ✓ Let volume, payload sensitivity, hops per request and time to detect choose the strategy, and record the choice and the rejected options in a decision record.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAR-P questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Integration alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.