Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 4 › Lesson 4.6

CCAR-P · Domain 4 · 16% of the exam · Lesson 4.6 · 23 min read

Monitoring a live Claude system: logs, dashboards, alerts and incidents

What each Claude call should log, which dashboards and alerts catch a quiet regression after a prompt change, and how to go from an alert to the failing step.

Written against objective 4.6 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

4.6.1 Why the wrong bin day took three weeks to surface

The City of Wrenmouth runs a 311 service, the non-emergency line for everything the council does, and Claude now answers most of it. On the website and by SMS, the assistant handles about 3,000 conversations a day: which day the bins go out, how to apply for a parking permit, where to report a pothole. For collection days the prompt tells the model to request a tool, get_collection_day, which the application runs against the waste-management system for the resident's address.

In May the digital team shipped prompt version 19. To make replies faster, a content editor had added a quick-reference table of collection days for the city's nine districts, copied from last year's leaflet. The model started answering from the table instead of requesting the lookup. For eight districts the table was right. Ashgrove's collection had moved from Tuesday to Thursday in the spring route change, so for nineteen days every Ashgrove resident who asked was told Tuesday. Olumide, head of 311 services, heard about it from two councillors and the local paper.

Green dashboards, wrong bin days

What the dashboards showed

99.7% of requests succeeded
Average reply 0.4 s faster
Spend on budget

What Ashgrove residents got

"Your bins go out on Tuesday"
Collection is on Thursday
19 days before anyone noticed
Every panel Wrenmouth watched looked normal or better after the change, while one district was being told the wrong collection day.

Tamsin, the architect Olumide asks to make sure it cannot happen again, starts with an uncomfortable finding: nothing Wrenmouth recorded could have revealed it. The logs held a status code and a duration per request. They did not say which prompt version answered, whether the lookup ran or which district asked, and nobody measured whether answers were right. The failure was not hidden. It was never written down.

Monitoring a Claude system is four connected pieces of work. A structured log record for every model call carries the versions and outcomes that produced it. Dashboards cut those records by segment. Alerts compare the numbers with a baseline and arrive with a runbook, the written first steps for whoever is paged. And the team practises following an alert down to the failing step.

4.6.2 The record every model call writes

Here is the question an incident asks: what produced this answer, what did it cost, and how did it end? Wrenmouth's log lines could only say whether a request succeeded. A record answers that only if it is structured: named fields a query can filter and count, not a sentence to read. Your code writes one for every model call, and a shared trace id, one identifier passed through every step of a request, joins the model calls and tool runs of one turn.

Field group What Tamsin records Why it earns its place
Identity Trace id, conversation id, Anthropic's request-id Joins the steps of a turn; gives support a handle on one call
Versions Model from the response, prompt version, tool and index versions, app release Ties every answer to the change that produced it
Segment Channel, intent, district, a pseudonymous resident key Lets the numbers be cut where residents live
Cost Input, output, cache-read and cache-write tokens Cost per conversation and cache health
Timing Latency per call, time to first text on streamed replies Percentiles per channel
Outcome stop_reason, tools requested, HTTP status and error type, results of your own checks The behaviour and error panels

Three fields need a word. stop_reason is the API's own note of why the reply ended: finished, waiting for a tool, cut off at the max_tokens cap, or declined. Take the model id from the response rather than your configuration, so the record shows what actually served even when a flag or fallback switched models. And record the cache fields. With prompt caching, the API reuses an unchanged prompt prefix, such as Wrenmouth's long system prompt, at a fraction of the input price, and input_tokens counts only the tokens after the cached part. Without the cache-read and cache-write counts, most of each prompt vanishes from your cost figures.

What stays out matters as much. A resident's message holds an address, a phone number by SMS, and sometimes more, such as a disability mentioned in a request for assisted collection. Tamsin logs the district, not the address; a keyed hash of the phone number, not the number; the intent and message length, not the text. Content that people must read to diagnose goes redacted into a separate store with named readers and short retention. API keys and authorisation headers never reach a log.

In this sketch, look at model and the prompt_version in ctx (what ran), tools (whether the model reached for the system of record) and request_id (a handle for Anthropic support). The setup of client, log and the constants is left out.

def answer(messages, ctx):  # ctx: trace_id, channel, district, intent, prompt_version
    start = time.monotonic()
    raw = client.messages.with_raw_response.create(  # raw gives access to the headers
        model=MODEL, max_tokens=800, tools=TOOLS,
        system=PROMPTS[ctx["prompt_version"]], messages=messages)
    msg = raw.parse()
    log.info(json.dumps({**ctx,                       # segment fields; no address, no phone number
        "model": msg.model, "request_id": raw.headers.get("request-id"),  # what ran; support handle
        "input_tokens": msg.usage.input_tokens, "output_tokens": msg.usage.output_tokens,
        "cache_read": msg.usage.cache_read_input_tokens or 0,
        "cache_write": msg.usage.cache_creation_input_tokens or 0,
        "latency_ms": round((time.monotonic() - start) * 1000),
        "stop_reason": msg.stop_reason,
        "tools": [b.name for b in msg.content if b.type == "tool_use"],  # lookup requested?
        "itpm_remaining": raw.headers.get("anthropic-ratelimit-input-tokens-remaining"),
    }))
    return msg

The failure path writes the same record with the HTTP status and error type in place of usage. The official SDKs retry 429s and server errors twice by default, so a record written around the call sees only the final outcome. To see every attempt, set max_retries=0 and retry in your own code, writing a record per attempt.

4.6.3 Dashboards that cut the numbers where users live

Why did Wrenmouth's dashboards stay green? They answered one question, "is the service up?", for the whole city at once. A useful dashboard asks five kinds of question (health, cost, behaviour, quality and feedback) and cuts every answer by channel, intent, district and version.

Panel What it tells you What Wrenmouth watches
Traffic by channel, intent, district Demand is normal, or one segment spiked or vanished Collection questions at 7 am on collection days
Latency p50, p95, p99 per channel How long residents wait, tail included SMS replies at p95 against a 10-second target
Errors by type: 400, 429, 529, other 5xx, timeouts Your bug or spend limit (400), your rate limits or spend cap (429), Anthropic's side (529, 5xx) Each type on its own line
Cost: tokens per conversation, cache-read share Spend per conversation; whether the cache still works Cache reads falling after a prompt edit
Stop-reason mix Replies cut short (max_tokens) or declined (refusal) SMS replies cut off mid-sentence
Tool-call rate per intent Whether answers still come from systems of record Collection answers that called get_collection_day
Quality and feedback: sampled scores, grounding checks, thumbs-down, repeat contacts Whether answers are right, and what users say Accuracy and "that's wrong" replies per district

Memorise the five kinds of question; the rows are one city's answer. Latency appears as percentiles: the p95 is the time within which 95% of replies finish, while an average blends the fast web majority with the slow SMS tail. Two quality signals need defining. A grounding check is code that compares a fact in the answer with the system of record, here the day named in the reply against the waste system's day. It is cheap enough to run on every answer. A sampled score comes from a reviewer, or a grader model with a rubric, marking a small random share of each day's conversations.

Segments are where quiet failures live. Ashgrove sends about one question in nine, and collection days are about a third of its questions. A city-wide accuracy figure therefore dipped by less than its normal weekly swing.

One regression, two views

City-wide, all questions

Sampled accuracy 95% to 92%within the usual weekly swing
Looks like noise

Ashgrove, collection days

Sampled accuracy 97% to 3%from the day v19 shipped
Unmistakable
The city-wide average absorbed Ashgrove's failure into normal noise; the same records split by district and intent show it at once.

Last, every prompt, model, tool or index release draws a change marker, a vertical line labelled with its version, across every panel. On Tamsin's dashboards, v19's marker at 10:05 on a Monday sits exactly where the tool-call rate for collection answers falls from 97% to 12%.

4.6.4 Alerts that notice a change, with a runbook attached

Which of these numbers should wake someone at night? Page on too much and the on-call team learns to ignore the pager; page on too little and residents become your alerting system. So page only for harm to users or an early signal of it, and send the rest to a queue reviewed in working hours.

Alert type When it wins What it costs
Static threshold (error rate above 2% for 10 minutes) Hard failures with a known bad level Misses anything that stays under the line
Baseline comparison (against the same hour last week, per segment) Behaviour and quality metrics with a weekly rhythm Needs history; storms and holidays trip it
Change window (new version against old, per segment, for 48 hours after a release) Quiet regressions after a prompt, model or tool change Needs the version on every record and enough traffic per version
Budget and headroom (spend or rate-limit headroom against plan) Cost and capacity Usage and cost reports lag by minutes; live headroom needs the headers logged

The hardest case is the one that burned Wrenmouth: a release that makes answers worse without failing any request. Two sets of signals catch it. The first is a behavioural fingerprint: cheap numbers computed on every request that move whenever behaviour moves, such as the tool-call rate per intent, reply length, the stop-reason mix and the handoff rate. The second is quality measured directly: grounding checks on every answer and scores on a sample, split by version. A fingerprint shift proves the system changed; the quality signals say whether it hurt.

A model upgrade gets the same change window, and the model field shows which snapshot served every answer. Not every shift is a regression, though. When a storm doubles reports of fallen trees, the questions have changed, not the assistant: that is input drift, so split by intent before blaming a release. Slow drift in quality is the opposite problem, because a week-on-week baseline drifts with it. Keep a fixed floor as well, such as the accuracy the service committed to, so a month of small slides still raises an alert.

An alert without a runbook hands the on-call engineer a riddle at 3 am. Here is Tamsin's runbook for the alert that would have caught v19.

Alert: collection answers without a lookup. Pages the 311 on-call during service hours.
Fires when: in any 30-minute window, fewer than 90% of collection-day answers in any district or prompt version called get_collection_day (baseline 97%).
What it means: bin days are coming from somewhere other than the waste system, such as a prompt table, an earlier turn or a guess. Requests still succeed, so nothing else will fire.
First checks: the collection panel split by prompt version and district; change markers in the last 24 hours; twenty records from the affected segment, opened by trace id in the restricted store.
Mitigate: set the prompt-version flag back to the last good version; no code deploy needed. If the waste system itself is down, switch collection questions to the "check the city website" reply.
Verify: lookup rate back above 95% within an hour and grounding-check failures at baseline. Owner: 311 digital team. Escalate to the waste services duty manager if residents were given wrong days.

4.6.5 What Anthropic's side can tell you

"Doesn't Anthropic already monitor all this?" It reports what your organisation sent and what it cost, not what your application did. Four sources are worth wiring in, each with a clear limit.

The Claude Console's Usage page breaks tokens down by model, workspace and API key, to the hour or minute. It lists requests blocked by rate limits and charts rate-limit use beside your cache rate, and its Cost page shows daily spend by workspace and model.

The Usage and Cost Admin API returns much the same data to your own systems. It needs an Admin API key, not an ordinary workspace key, and reports usage in 1-minute, hourly or daily buckets and cost in daily buckets, typically within five minutes. Tamsin gives each application its own workspace, the Console's container for API keys, which can carry its own spend and rate limits. Finance then sees spend per service, and the team reconciles its own token counts with the bill.

Response headers work per call. Every response carries a request-id, and the anthropic-ratelimit-* headers give the limit, remaining capacity and reset time for requests, input tokens and output tokens. Log the remaining input tokens and you can alert on headroom before 429s start.

The errors themselves need telling apart, because each has its own runbook. A rate-limit 429 carries a retry-after header. The 429 for your usage tier's monthly spend cap has none, carries the error code enforced_spend_limit_reached, and fails on every retry until access resumes. A spend limit you set yourself, on the organisation or a workspace, returns a 400 instead, so a wall of 400s is not always a bug. A 529 means the API is temporarily overloaded across users. And a sharp jump in your own traffic, such as the morning after a storm, can trigger 429s from acceleration limits.

Anthropic's agent tools export OpenTelemetry, the open standard for metrics, logs and traces, to a collector you choose. Claude Code, Anthropic's agentic coding tool, starts exporting once you set CLAUDE_CODE_ENABLE_TELEMETRY=1 and an exporter. It sends token and cost metrics, plus an event per API request (model, cost, duration, tokens, request id) and per failed request (status, attempt count). Prompt and response text stay out unless you opt in, and traces are in beta. The Agent SDK, the library for running the same agent loop inside your own application, starts the Claude Code CLI underneath and exports the same data. Wrenmouth's overnight agent, which turns pothole reports into work orders, exports to the 311 collector with OTEL_SERVICE_NAME set, so its records are easy to filter.

Source Answers Cannot answer
Your own structured records What ran, for which segment, and how each answer ended The official bill
Console pages and the Usage and Cost Admin API Tokens by workspace, model and key; spend by workspace and model; rate-limit use and cache rate Anything about one request or prompt version
Rate-limit headers and request-id Live headroom on each call; a handle for support Why an answer was wrong
Claude Code and Agent SDK telemetry Agent activity: requests, tokens, errors, tool results Your web application's own steps; the bill, since its costs are estimates

4.6.6 From alert to failing step

The pager goes off. It is tempting to start fixing at once: rewrite the prompt, try a bigger model, read transcripts at random. Resist it. The records already hold the path from symptom to cause.

From alert to failing step

ALERTa metric leaves its baseline
SCOPEwhich channel, segment, version
ALIGNwhich change marker
INSPECTsampled records, step by step
MITIGATEroll back that change
VERIFYthe same panel recovers
Each step narrows the search with a field the records already carry, so the team reaches the failing step and a safe rollback in minutes.

Replay May with Tamsin's design. Prompt v19 goes live at 10:05, and at 10:35 the lookup alert fires: 12% of collection answers called get_collection_day, against a baseline of 97%. Every district and both channels are affected, and the drop starts at v19, the only change marker in the window. The on-call engineer opens twenty v19 collection answers with no tool call, by trace id, in the restricted store. Every one quotes the new table, and their grounding checks show eight districts right and Ashgrove wrong. That is the failing step: the answer skipped the lookup and read a stale table.

Mitigation comes before the root-cause fix. The engineer sets the prompt-version flag back to v18: no deploy, and reversible. Within the hour the lookup rate is back at 97%. The repair to v19 follows as a normal change, and the Ashgrove question becomes a permanent test case. A restricted lookup of the pseudonymous keys lets the service text a correction to the few residents told Tuesday. The harm lasted about an hour, not nineteen days. Deciding why a step failed is diagnosis; monitoring gets you to the step fast.

4.6.7 The exam traps

Every trap here either watches the wrong thing or reacts before looking.

  • ✗ Logging a line of free text, or the full conversation, per request. ✓ Write one structured record per call with versions, tokens, stop reason, tools and outcome. Keep raw personal data out and any content redacted and restricted.
  • ✗ Watching city-wide averages. ✓ Show latency as percentiles and cut every panel by segment and version. A failure in one district or intent hides inside the average.
  • ✗ Alerting only on errors and latency. ✓ Add behaviour and quality signals with a change window after every release. A quiet regression returns HTTP 200 on time.
  • ✗ Treating every error alike, or raising retries until errors disappear. ✓ Split 400, 429, 529 and other 5xx, log the rate-limit headers, and give each its own runbook. More retries hide the signal and add load.
  • ✗ Explaining one bad answer from the Console or the Usage and Cost Admin API. ✓ Use them for spend, capacity and reconciliation. They aggregate by workspace, model and key; per-request questions need your own records.
  • ✗ Answering an alert by switching to a bigger model or rewriting the prompt. ✓ Follow the alert to the failing step, roll back the change it aligns with, verify, then fix. The cause is often the pipeline, not the model.

4.6.8 Put it together: instrument an assistant and catch a quiet change

You now have every piece, from the record each call writes to the path from alert to failing step. The fastest way to own them is to watch a quiet change slip past a health check and get caught by a fingerprint.

From here the loop closes. Every incident becomes a test case in your offline evaluation set (4.2), so the same failure is caught before the next release. Diagnosing system issues (4.4) picks up where the incident flow stops, deciding whether the failing step was a prompt failure, a hallucination or a model mismatch. SLAs (6.3) are reported from these same dashboards, and Claude Code can support the on-call investigation itself (7.3).

Key takeaways

  • ✓ Quiet changes to an LLM system shift behaviour without raising errors, so uptime metrics alone cannot see them.
  • ✓ Your code writes one structured record per model call with identity, versions, segment, tokens including cache, latency, stop reason, tools and outcome, and no raw personal data.
  • ✓ Dashboards cover health, cost, behaviour, quality and feedback, show latency as percentiles, carry change markers, and cut every panel by segment and version.
  • ✓ Alerts page on user harm or its early signal, compare against baselines and a change window after every release, and each comes with an owner and a runbook.
  • ✓ The Console and the Usage and Cost Admin API report spend, tokens and rate-limit use by workspace and model; headers give live headroom; Claude Code and the Agent SDK export OpenTelemetry; none explains a single answer.
  • ✓ An incident runs alert, scope, align, inspect, mitigate by rolling back, and verify, with the root-cause fix afterwards as a tested change.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

30 CCAR-P questions on Domain 4, free

Every question in the bank is tagged to a domain, so you can drill 30 questions on Evaluation, Testing & Optimization alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 4 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources