Home › Study guides › CCAR-P › Domain 6 › Lesson 6.3
CCAR-P · Domain 6 · 14% of the exam · Lesson 6.3 · 22 min read
Expectations, SLAs and feedback loops for an LLM system
How to replace a promise of perfection with measured expectations, write SLAs that cover quality and the provider, and run feedback loops that close.
Written against objective 6.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
6.3.1 Why "no mistakes" was never a requirement
Latchmere Security, a cybersecurity software company, launched a contract-review assistant for its legal department four months ago. Sixty lawyers upload customer and supplier contracts, and Claude compares each clause with the department's playbook, its standard positions on liability caps, indemnities, data processing and governing law. It flags each deviation it finds with the clause number and suggested fallback wording, and a lawyer decides. The assistant works, yet Rosamund, the solution architect who designed it, has three problems on her desk.
Gideon, the general counsel, approved the project believing it would make "no mistakes", and a missed indemnity clause last month shook his confidence in the whole tool. Lawyers email feedback to a shared inbox nobody owns, where about 140 messages sit unread. And during an acquisition closing, the assistant returned errors for three hours while the deal team worked through the night. The partner leading the deal complained to Gideon, and nobody could say what had happened.
None of these is a model problem. Each is a missing agreement: on quality and how it is measured, on the availability lawyers can count on and what happens when it fails, and on where feedback goes and who answers it. A large language model (LLM) gives probabilistic answers, so some fraction of reviews will contain an error however good the design. Expectation alignment is the architect's work of turning hopes into explicit, measured agreements that everyone can check.
From assumptions to agreements
Assumed
Agreed
6.3.2 Set expectations before launch, in measured terms
Here is the belief that causes most of the trouble. A sponsor who hears "the assistant reviews contracts" hears "correctly, every time". It is tempting to leave that belief alone, because it helped get the project approved. Resist it: a belief nobody corrected becomes a promise you made, and the first visible error breaks it.
Think of an honest weather forecaster: "seventy percent chance of rain" earns more trust than a promise of sunshine, because the number can be checked. An expectations statement works the same way: what the system does, its measured errors on your own data, its limits, and who stays accountable. Anthropic's guidance on success criteria asks for this shape: specific, measurable, achievable and relevant, usually across several criteria rather than one accuracy number.
Break the error rate down by type, because the types cost different amounts. At Latchmere a missed deviation, a non-standard clause left unflagged, is the expensive error, because a lawyer may trust the silence. A false flag costs a minute of reading. One blended "92% accurate" hides that difference. Measure on your own data: two senior lawyers labelled every deviation in 300 past Latchmere contracts, and those contracts became the evaluation set, the fixed cases every later change is tested against.
Latchmere skipped this step at launch, so Rosamund does it now, as a reset. Here is the statement she agreed with Gideon in place of "no mistakes"; look at the second line, which gives the measured misses and what the lawyer must still do.
What it does: compares each clause of an inbound customer or supplier contract with the legal playbook and flags deviations with the clause number and suggested fallback wording. Contracts in English under English or New York law, up to 80 pages.
How often it is wrong (measured on 300 past contracts reviewed by two senior lawyers): it missed 4 in 100 playbook deviations, and 12 in 100 of its flags were false. Misses cluster in clauses split across schedules. A lawyer reads every contract in full; a clean review means "no deviation found", never "no deviation".
What it will not do: give a legal opinion, approve or sign anything, read scanned handwritten amendments, or assess any other governing law.
Who is accountable: the reviewing lawyer for every decision; the architecture team for the measured rates, re-measured monthly and reported to the general counsel.
The conversation changes. Gideon no longer asks whether the assistant made a mistake, but whether the miss rate still holds at 4 in 100 and whether lawyers still read in full. The next missed clause becomes a data point to check against the agreed rate, not a betrayal.
6.3.3 SLOs, SLAs and error budgets for a probabilistic system
Once the rates exist, the question is which of them the service commits to, and what happens when one slips. The vocabulary comes from site reliability engineering. A service level indicator (SLI) is what you measure. A service level objective (SLO) is the target the team runs to. A service level agreement (SLA) is the commitment to the stakeholder, with a stated consequence when it is missed. For an internal service that consequence is rarely money: it is an incident review, a report to the general counsel and priority for the fix. Set each SLA line a little looser than its SLO, so the team sees trouble before the promise breaks.
An LLM system needs four kinds of line; teams most often leave out the third. Latency lines use a percentile: the p95 is the time within which 95% of reviews finish.
| Line | How to measure it | Latchmere's SLA line (the team's SLO in brackets) |
|---|---|---|
| Availability | A synthetic probe: a short test contract every five minutes in business hours and any declared closing, counted only when a complete review returns | 99% of probe runs per month (99.5%) |
| Latency | A percentile, end to end from upload to finished review | p95 under 3 minutes for contracts up to 80 pages, per month (2.5 minutes) |
| Quality | The evaluation set before every release; a weekly random sample of 40 live reviews checked by two senior lawyers | Missed deviations at most 5 in 100 over the last four weekly samples (4 in 100) |
| Escalation | Time from a report to acknowledgement, to degraded mode and between updates | Severity 1 in a declared closing: acknowledged in 15 minutes, updates every 30 (10 and 20) |
Gideon's first ask was 99.9% around the clock. Negotiating an SLA means pricing each line in front of the person who wants it. Round-the-clock 99.9% would have needed a night on-call rota and a second route to the models, for a tool lawyers use mostly in office hours. They settled on business hours plus any declared closing, which covers the nights that matter, at a target the redesigned service can meet. Every line was agreed the same way. Start from what the stakeholder's work needs, show the measured baseline and the cost of each step up, and commit only to what you can measure and the dependency supports.
Three details make these lines fit an LLM. Measure from the user's side: a response cut off at its token limit returns successfully with half a review, and the lawyer counts that as a failure. Use both quality instruments, because the evaluation set covers the contracts you thought of and the weekly sample catches the rest. And state every quality number as a rate over a window, with its sample size.
An error budget turns an SLO into a decision rule. It is the unreliability the SLO allows: 100% minus the target. Latchmere's business hours come to about 308 hours a month (14 hours a day, 22 working days), so a 99.5% availability SLO allows about 92 minutes of failed probes. Treat it as a spending allowance for risk. While budget remains, prompt changes and upgrades ship at the normal pace; once it is spent, the agreed policy freezes everything except reliability fixes. The same rule covers quality: if the four-week miss rate passes 4 in 100, changes stop until the cause is found. The three-hour closing outage would have spent almost two months of budget in one night.
6.3.4 The provider is inside your SLA
The post-incident review found two causes for the closing outage. Most failed requests were 429 errors, because a job re-extracting clauses from 18,000 archived contracts had used up the organisation's rate limit. It ran on the same model, in the same workspace (a group of API keys that can carry its own limits). For about forty minutes, Anthropic's status page also showed elevated API errors, and some requests failed with 529, the overloaded error. Anthropic's SDK (the official client library) retried each failure twice, its default, then gave up, and nothing in the design said what came next.
Your SLA sits on a dependency you do not control, so design for its documented behaviour. Rate limits apply per organisation and per model (a few older models share one combined limit), and the docs call them maximums, not guaranteed minimums. A sharp jump in usage can trigger 429s from acceleration limits. Hitting the monthly spend cap also returns a 429, without a retry-after header, and no retry succeeds until access resumes. A 529 means the API itself is temporarily overloaded, which can happen under high traffic across all users.
The standard service tier, the default for every request, has best-effort availability. Priority Tier, which prioritises requests and targets 99.5% uptime, is no longer sold: existing commitments run until their contracts end, and for guaranteed capacity the docs point to Anthropic's sales team. Incidents are posted per component, the Claude API among them, at status.claude.com, where you can subscribe to updates. A restaurant that promises dinner by eight plans for the evening the supplier's van is late. Rosamund's redesign plans for the provider's bad evenings, layer by layer.
Where a review goes when the provider falls short
| Option | When it wins | What it costs |
|---|---|---|
| Isolate bulk work: its own workspace with a capped rate limit, or the Message Batches API, which has separate limits | Background jobs compete with people for the organisation's limits | Batch results arrive asynchronously, so only work that can wait moves there |
| Retry with backoff | Short bursts of 429 or 529 errors | Adds latency; useless against the spend cap; unbounded retries add load |
| A fallback model | One model is at its limit, and the fallback has a limit of its own | It must pass the same evaluation set, and its quality may differ |
| A second route to Claude through Amazon Bedrock, Google Cloud or Microsoft Foundry | The availability target is beyond what one route can support | A second contract, quotas and data-handling review; model lists differ; every change tested twice |
| A degraded mode | Always, as the floor when everything else fails | Slower reviews and lawyers' time |
Latchmere adopted every option except the second route, because its data-processing terms covered the Claude API only. A second route also helps only if it does not share the failure: AWS and Google run the inference behind Bedrock and Google Cloud, while Anthropic itself operates the Claude service in Microsoft Foundry. Rosamund also added a closing-window protocol: the deal team declares a closing 48 hours ahead, bulk jobs pause, the on-call engineer checks rate-limit headroom, and Severity 1 times apply. The SLA states the dependency and counts provider incidents against the error budget. The lawyers cannot tell whose outage it was, and the architect chose the dependency.
6.3.5 Feedback loops that close
The 140 unread emails were not a shortage of feedback but a loop with no owner, no structure and no way back to the sender. Lawyers who write in and hear nothing stop writing, and the next problem reaches the general counsel as a complaint instead of reaching the team as a report.
A working loop draws on four inputs. In-product feedback is a button on each flag and review, with a reason code (missed deviation, false flag, wrong citation, too slow, other) and free text. It also records the review's trace ID, so the team can reproduce the case. Reviewer notes are the signal lawyers give without meaning to: every dismissed flag and every deviation added by hand, captured as data. The weekly quality sample is the third input. The fourth is a monthly service review with Gideon and the practice leads, where the SLO report, the main themes and scope requests get decided.
The feedback loop at Latchmere
Triage gives each item a category, because categories have different owners. An item can be a prompt or retrieval defect, a gap in the evaluation set, a playbook content error for the legal team, a scope request, a training need or a provider incident. Priority comes from severity and frequency, so one missed liability cap outranks thirty complaints about formatting. Every fixed defect also becomes a case in the evaluation set, so the same mistake cannot return unnoticed with the next prompt change.
Closing the loop is the step most teams skip. Each person who reported something gets a status: fixed in which release, planned for when, or declined and why. A monthly "you said, we did" note tells the department what changed. Lawyers who see their reports acted on keep reporting, and the team keeps its best source of production evidence.
6.3.6 Communicating changes and holding scope
Two things erode a well-aligned service after launch: changes users did not see coming, and scope that grows until the measured rates no longer describe the system. To a lawyer, a new model that flags more readily looks like the tool got worse overnight, even when its miss rate improved.
Each Claude model ID names a fixed snapshot, and an updated model ships under a new ID, so a model change is always yours to make and announce. You cannot stay on one ID forever, because each has its own deprecation and retirement schedule. The serving infrastructure around a fixed model can also change, occasionally shifting behaviour slightly on the same ID, so the weekly sample never stops. Rosamund announces every change before it ships and outside any declared closing. Look at the third line: before-and-after rates by error type, and what the lawyers will notice.
Change: the contract-review assistant moves to a newer Claude model on 14 October, first for the commercial team (12 lawyers) for one week, then for everyone.
Why: the current model is deprecated and retires in January.
Evaluation on the 300-contract set: missed deviations 4 in 100 before, 3 in 100 after; false flags 12 in 100 before, 17 in 100 after; p95 time unchanged. You will see more flags on limitation-of-liability clauses, and most of the extra flags are borderline and quick to dismiss.
Rollback: one configuration switch returns the current model within an hour. Report anything unexpected with the in-product button, reason "other".
Scope creep arrives as reasonable requests: could it also draft the reply to the counterparty, review German-law contracts, or say whether to sign? Each is a new capability with no measured rate, and switching it on quietly puts unmeasured behaviour under the SLA's name. Rosamund's rule is that a new capability enters the backlog as a scope request and launches only with its own evaluation set, expectations statement and SLO lines, decided at the monthly review. "Not yet, and here is what it would take" is an answer partners can plan around.
6.3.7 The exam traps
Every trap below leaves an expectation implicit or unmeasured. The fix is an explicit agreement with a number, an owner and a mechanism that keeps it true.
- ✗ Letting a sponsor's "no mistakes" stand, or answering it with one blended accuracy figure. ✓ Agree measured rates by error type on your own data, what the system will not do and who stays accountable, before launch.
- ✗ An SLA of uptime and average latency only. ✓ Add quality measured on evaluations and sampled review, latency as a percentile, and escalation times.
- ✗ Measuring availability as the model API's success rate. ✓ Measure end to end from the user's side, count only complete results, and count provider incidents against the error budget.
- ✗ Promising more availability than the dependency supports, or planning to buy Priority Tier. ✓ Treat the standard tier as best-effort and rate limits as maximums, isolate bulk work, and keep a tested fallback and a degraded mode.
- ✗ Collecting feedback in an unowned inbox, or feeding every complaint straight into the prompt. ✓ Triage weekly under an owner into a prioritised backlog, turn fixes into eval cases, ship through tested releases, and tell each submitter what happened.
- ✗ Shipping a model or prompt change silently, or switching on a requested capability without measuring it. ✓ Announce changes with before-and-after results and a rollback, and give each new capability its own evaluation, expectations and SLO.
6.3.8 Put it together: write the service agreement and test it
You now have every piece, from measured expectations to rules for change and scope. The quickest way to make them stick is to measure a small version and watch a careless change break a promise you wrote down.
These agreements feed the rest of the domain. Documenting architectures (6.4) writes the SLA, the degraded mode and the closing-window protocol down for the team that will run them. Supporting lifecycle phases (6.5) makes the monthly service review part of monitoring and iteration after handoff. Operational issue resolution (7.3) is what happens inside the Severity 1 you defined here.
Key takeaways
- ✓ An LLM system cannot promise "no mistakes"; agree before launch on measured error rates by type, what it will not do and who stays accountable.
- ✓ An SLI is what you measure, an SLO is the team's target, and an SLA is the commitment with a consequence, negotiated from the measured baseline and set a little looser.
- ✓ An LLM SLA covers availability, latency percentiles, quality measured on evaluations and sampled review, and escalation times, all from the user's side, with error budgets that decide when to ship and when to fix.
- ✓ The provider sits inside your SLA: rate limits are maximums, the standard tier is best-effort and Priority Tier is no longer sold, so isolate bulk work, bound your retries, and keep a tested fallback and a degraded mode.
- ✓ A feedback loop captures structured signals, triages them under one owner into a prioritised backlog, turns fixes into evaluation cases, and tells each submitter what happened.
- ✓ Every prompt, playbook or model change is announced with before-and-after evaluation results and a rollback, and every new capability needs its own evaluation and SLO before launch.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
27 CCAR-P questions on Domain 6, free
Every question in the bank is tagged to a domain, so you can drill 27 questions on Stakeholder Communication & Lifecycle Management alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 6 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.