Home › Study guides › CCAR-P › Domain 4 › Lesson 4.3
CCAR-P · Domain 4 · 16% of the exam · Lesson 4.3 · 21 min read
A/B testing and iterative improvement: proving a change on live traffic
How to improve a Claude feature one tested change at a time: error analysis, offline evals, A/B test design, shadow and canary releases, and reading results.
Written against objective 4.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.3.1 Why a prompt that reads better is not yet a better product
Nestmark is a property-listing portal. When an estate agent adds a home, they enter its facts (rooms, floor area, features, photos) and Claude drafts the description. The agent accepts the draft, edits it or throws it away before the listing goes live. About 12,000 listings a week pass through this step, from roughly 2,600 agency branches. Today it runs prompt v7 on Claude Sonnet 5.5.
Lorcan, the product lead for agent tools, has spent a week on prompt v8, and on the ten listings he tried it reads clearly better. Finance adds that Claude Haiku 4.5 costs half as much per token as Sonnet 5.5, so Lorcan plans to ship both on Monday. Maelle, the architect, points out that the prompt's author chose and judged those ten listings. Nestmark cares about three other things. Do agents publish the draft without rewriting it? Does it cost less? And does it avoid the complaint that hurts most, a buyer at a viewing who finds the "sea views" were invented?
A change that reads better can still lose on all three. Agents may keep editing out a new stock phrase, the cheaper model may invent more features on long listings, and nobody will know which change caused what. The discipline that prevents this is an improvement loop. Each change starts as a hypothesis drawn from real failures and passes an offline evaluation. Then it proves itself on live traffic against the current version, usually in an A/B test, and ships or reverts on rules agreed in advance.
4.3.2 The improvement loop: one hypothesis, one change
Where should the next change come from? The tempting source is a hunch. The reliable source is error analysis, reading the outputs that failed and sorting them by cause. Maelle samples 200 drafts that agents edited last month and compares each with the published text. She tags every edit with a reason: an invented feature (the "garden" that is a paved yard), a missing fact agents always add, a stock phrase such as "nestled", the wrong length. Invented features and stock phrases explain about 60% of edits.
That gives a testable hypothesis. Most edits remove invented features and stock phrases. So a prompt that names the supplied facts as the only source, and explains a house style instead of banning words, should raise the share published unedited. Prompt v8 is that change and nothing else.
One turn of the improvement loop
Lorcan's plan bundles two changes, a new prompt and a new model. If the bundle wins, you cannot tell whether the model quietly lost quality that the prompt hid. If it loses, you do not know which half to revert. Prompts and models also interact: Anthropic's cost guide measured prompts written for one Opus model costing 36% more per ticket on its successor, with no change in accuracy. So Maelle gives each change its own arm.
Each arm is a variant: a named, immutable bundle of prompt version, model ID and parameters, here the effort level, which trades thoroughness for tokens and time. Immutable means v7 is never edited in place; a fix becomes a new version. Every Claude model ID names a pinned snapshot, so the manifest records exactly what ran; older models' short aliases, such as claude-haiku-4-5, are only pointers. Look at the three arms, each one change from the one before, and at the last line, which splits every metric by arm.
experiment: listing-desc-2026-10 # also salts the account hash
unit: account_id # a branch keeps its arm for the whole test
arms:
A: # control: what runs today
prompt: listing-desc/v7
model: claude-sonnet-5-5
effort: low
B: # change 1: the prompt
prompt: listing-desc/v8
model: claude-sonnet-5-5
effort: low
C: # change 2: the model
prompt: listing-desc/v8
model: claude-haiku-4-5-20251001 # no effort parameter on Haiku 4.5
split: {A: 34, B: 33, C: 33}
log_with_every_draft: [experiment, arm, prompt, model] # split any metric, roll back by arm
The first gate is offline. Anthropic's model guide calls a good evaluation set, run with your actual prompts and data, the most important step in deciding whether to change models. Maelle runs all three variants over 400 past listings, three trials each because outputs vary between runs. Code checks cover length and required facts, and a grader flags unsupported claims; building that set is its own topic. The gate rule: no variant may invent more than A.
| Variant | Rubric pass rate | Unsupported claims | Cost per draft |
|---|---|---|---|
| A: v7 on Sonnet 5.5 (live) | 78% | 3.4% | baseline |
| B: v8 on Sonnet 5.5 | 87% | 0.8% | +4% |
| C: v8 on Haiku 4.5 | 85% | 1.2% | -52% |
Both candidates pass. The table cannot say whether agents will publish these drafts unedited; that is behaviour, and only live traffic shows it. Anthropic's evals guidance agrees: automated evals are the first line of defence on every change, and A/B testing validates significant changes once traffic is sufficient.
4.3.3 Designing the online test: unit, split and metrics
Why not ship v8 to everyone and compare this month with last month? Because everything else changed too. The season shifts the mix of flats and family houses, agents' workloads change, and Anthropic notes that serving infrastructure updates can cause minor behaviour differences on an unchanged model ID. A concurrent control, the old variant running at the same time on comparable traffic, sees the same world. Any difference left is the change.
The first decision is the unit of randomisation, the thing you assign to an arm. Assigning each request is the easy default, and it is wrong here. One branch would see v7 on one listing and v8 on the next. Agents who catch an invented garden in a v7 draft start checking every draft, v8's included, and the arms blur together.
It is like trialling two teaching methods by switching method every lesson: each class learns from both, so no class shows what either method does alone. The outcome is a habit of the agent, so the account is the unit. Your code hashes the account ID with the experiment name, so each branch keeps its arm for the whole test and the next experiment splits branches afresh.
Randomise by request or by account
Per request
Per account
The general rule: randomise at the level where the outcome lives and where the experience must stay consistent. For a conversational feature that is the user or the session, never the request, because a conversation whose prompt changes halfway is neither variant. Analyse at the same level, too: a busy branch's 60 drafts reflect one office's habits, not 60 independent observations.
One primary metric decides the test: the share of drafts published without edits, counted per account, a form of what Anthropic's success-criteria guide calls implicit user feedback. Defining it precisely is a metrics question of its own. Guardrail metrics carry thresholds that stop an arm whatever the primary says. Nestmark's are cost per published description, p95 generation time (the time 95% of drafts beat), buyer reports of inaccurate listings and the graded unsupported-claim rate. A win that breaches a guardrail is not a win.
Four inputs decide how long the test runs: the baseline rate (42% published unedited), the smallest lift worth shipping (3 points), the significance level (how often you accept a false win) and the power (how reliably you catch a real one). With the usual 5% and 80%, and independent drafts, a 3-point lift needs about 4,300 drafts per arm. But the same agents edit all of a branch's drafts, so each extra draft adds less information. At Nestmark's branch sizes that roughly triples the requirement, and at about 4,000 drafts per arm per week the test runs four full weeks.
Arm C, a cost saving, does not need to win. It must show it is no more than 3 points worse than B, a non-inferiority margin agreed now. Because B must beat A by 3 points to ship, a C within the margin is still no worse than today's service, at half the price. A 3-point margin needs about the same sample as a 3-point lift. The traffic split trades exposure for time: equal arms reach the sample fastest, while a candidate on 10% of traffic would need almost three times as long. Maelle limits exposure earlier, with shadow and canary stages.
Maelle writes all of it down before any traffic moves, so nobody can choose the metric after seeing the numbers.
Hypothesis: most agent edits remove invented features and stock phrases; prompt v8 (supplied facts as the only source, house style with its reason) raises the share of drafts published unedited. Claude Haiku 4.5 runs v8 no worse than Claude Sonnet 5.5 at half the per-token price.
Arms and unit: A v7 on Sonnet 5.5 (control), B v8 on Sonnet 5.5, C v8 on Haiku 4.5; assigned by agency branch account, hashed with the experiment name; split 34/33/33 after a week of C in shadow and a three-day canary of C on 5% of accounts.
Decision rule: B ships if its lift over A is statistically significant and at least 3 points; C replaces B if it is no more than 3 points below it. One read, at the end.
Guardrails (stop a candidate arm if breached, watched daily): cost per published description more than 10% above A's; p95 generation time over 8 seconds; buyer inaccuracy reports above 1.5 per 1,000 listings; unsupported claims above 2% of graded samples.
Duration: four full weeks. Segments declared now: price band (premium = top tenth), property type, branch size.
Rollback: map every account back to arm A; a configuration change, no deploy.
4.3.4 Shadow, canary, A/B or side-by-side: choosing the test
An A/B test is the strongest evidence you can get, but it is not always possible and rarely comes first. Nestmark's commercial-property section publishes about 70 listings a month; detecting a 3-point lift there would take years. Each alternative answers a narrower question at a lower price.
| Option | When it wins | What it costs |
|---|---|---|
| Offline eval | Every change, as the first gate; reproducible, no user impact | Only as good as the eval set; cannot measure user behaviour |
| Shadow mode: the candidate runs on live inputs, its output logged but never shown | Checking cost, latency, failures and graded quality on real traffic with zero exposure | You pay for both variants' calls; no signal on what users do |
| Canary release: a small share of accounts gets the candidate, with automatic rollback on guardrails | Limiting the blast radius of a risky change | Too small to measure the primary metric; a safety step, not a verdict |
| A/B test | Enough traffic and a user outcome to measure | Weeks of runtime; tests only what you deploy; shows what changed, not why |
| Side-by-side preference review: experts judge blind pairs of old and new outputs | Low traffic, or quality that only a specialist can judge | Slow and costly; raters disagree; measures preference, not behaviour |
These options are stages, not rivals. Maelle runs arm C in shadow for a week on every new listing. The offline set held few photo-heavy luxury listings, and she wants Haiku's p95 time and failure rate on the real mix. A three-day canary on 5% of accounts follows, watched on the guardrails, and only then the full test. Anthropic's cost guide ends its measurement method the same way: shadow the winner on a traffic slice before cutover.
How a change earns full traffic
For the commercial section, Maelle replaces the A/B test with a side-by-side review. Three senior commercial agents see 80 past listings, each with the v7 and v8 drafts in random left-right order and no labels, and pick the one they would publish. Paired, blind judgements detect a real difference with far fewer cases than an A/B test needs, because each pair cancels out the listing's own difficulty. The majority preferred v8 on 57 of the 80 listings, well beyond chance, so the commercial section gets v8 too.
4.3.5 Reading the result: long enough, large enough, for whom
At the end of week one, arm B leads A by 6 points with p below 0.05, and Lorcan wants to ship. (The p-value is the chance of a gap at least this large if the arms were identical.) If the result is significant, why wait? Because a 5% false-positive rate holds only for ONE look at the planned end. Every extra look is another lottery ticket for a false win. In simulation, an A/A test (two identical arms) checked daily for a month declares a false winner in more than a quarter of runs. Stop at the first good day, and a lucky day ships.
Week one is also the least typical week. A novelty effect is a change in behaviour caused by newness rather than quality: agents curious about the new style accept more at first, while wary ones edit more. Weekly cycles matter too, since branches list heavily early in the week. Run whole weeks, and compare the first with the last to see whether an effect settles; Nestmark's settled lower.
After four weeks, each difference comes with its 95% confidence interval (CI), the range the true difference plausibly lies in.
| Comparison | Published without edits | Guardrails | Reading |
|---|---|---|---|
| B vs A (the prompt) | +3.6 points (CI +1.7 to +5.5); week one was +6.1 | Cost +4%, complaints flat | Real and above the 3-point bar: ship v8 |
| C vs B (the model), all listings | -0.7 points (CI -2.6 to +1.2) | Cost -54%, complaints flat | Within the 3-point margin |
| C vs B, premium listings (a tenth of volume) | -5.2 points (CI -8.9 to -1.5) | Agent support tickets up | Fails the margin in a declared segment |
Statistical significance says a difference is probably not noise; practical significance asks whether it is big enough to act on. With enough traffic, a 0.4-point lift can be statistically significant and still not worth a migration, which is why the plan fixed the smallest worthwhile effect in advance. B clears both bars. A saving is read the other way round: C passes overall because its whole interval stays above minus 3 points.
The premium row is a segment effect: a change that holds up overall but hurts one group. It counts as evidence here because price band was declared before the test and the gap held in every week. A segment you find afterwards, by slicing the data until something shows, is only a hypothesis for the next turn of the loop. Slice enough ways and one slice will look significant by chance.
Maelle's decision record says: ship v8 everywhere, and route standard listings to Haiku 4.5 and premium listings to Sonnet 5.5. That routing rule is itself a new variant, so it gets its own short canary. Shipping means moving traffic to the winning variant's ID while the old one stays deployable, so a rollback is a configuration change, not a deploy. Had B missed its bar, the revert would have been the same switch.
4.3.6 The exam traps
Most traps here make the result untrustworthy: the wrong comparison, the wrong unit, or a decision taken too early.
- ✗ Shipping because the new version looked better on a handful of examples or a public benchmark. ✓ Gate on an offline eval of your own data, then test user outcomes online. Hand-picked cases judged by their author are anecdotes.
- ✗ Comparing this month on the new version with last month on the old. ✓ Run a concurrent control on comparable traffic. Before-and-after results mix the change with season, traffic mix and every other shift.
- ✗ Randomising each request, or letting users opt in to a beta. ✓ Assign a stable unit (user, session or account) at random, and analyse at that unit. Per-request assignment contaminates conversational features; volunteers differ from everyone else.
- ✗ Testing a new prompt and a new model together in one arm. ✓ Keep one change per comparison, with extra arms or sequential tests. A bundle cannot be attributed or partly reverted.
- ✗ Stopping on the first day the result is significant, or choosing the metric after seeing the data. ✓ Fix the primary metric, smallest worthwhile effect and duration in advance, and read once at the end. Only a guardrail breach stops a test early.
- ✗ Declaring a winner on the primary metric while cost or complaints regressed. ✓ Treat guardrails as vetoes. A significant lift below the worthwhile effect, or bought with a guardrail breach, does not ship.
4.3.7 Put it together: plan a test and watch it fail
You now have the whole loop, from error analysis to an online test with the right unit, metrics and duration. The two pieces that most often go wrong are assignment and stopping, so build both and watch each one fail.
The rest of the domain feeds and follows this loop. Diagnosing system issues (4.4) deepens the error analysis that starts each turn. Optimising token usage, latency and cost (4.5) supplies many of the changes the loop tests. Monitoring (4.6) watches the shipped variant after the test ends and hands the loop its next failures.
Key takeaways
- ✓ A change is a hypothesis from error analysis: one cause, one change, one versioned variant with its prompt, pinned model ID and parameters logged on every output.
- ✓ An offline eval gates every change, but only live traffic measures what users do with the output.
- ✓ A/B tests need a concurrent control and a stable unit of randomisation (account, user or session, never the request for conversational features), analysed at that unit.
- ✓ Fix the primary metric, guardrail thresholds, smallest worthwhile effect or non-inferiority margin, and duration before the test starts; an equal split reaches the sample fastest.
- ✓ Shadow mode, canary releases and blind side-by-side review answer narrower questions with less exposure or less traffic, and usually come before or instead of an A/B test.
- ✓ Read the result once at the planned end, after full weekly cycles; ship a lift only if it is statistically and practically significant with guardrails intact, and keep rollback a configuration switch.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
30 CCAR-P questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 30 questions on Evaluation, Testing & Optimization alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.