Home › Study guides › CCAR-P › Domain 4
CCAR-P · Domain 4 of 7 · 6 lessons · about 133 min
Domain 4: Evaluation, Testing & Optimization
Measuring and improving a Claude solution: metrics, evaluation datasets, A/B tests, diagnosing failures, cost and latency, and monitoring. 16% of CCAR-P.
This domain is 16% of CCAR-P, about 10 of the 63 questions. It covers how you know a Claude solution works, how you find out why it stopped working, and how you make it cheaper and faster without making it worse.
The vocabulary: an evaluation (eval) runs a fixed set of test cases through the system and scores the results. Code-based grading checks answers mechanically; LLM-as-judge has a model score answers against a rubric; human evaluation is the reference both are checked against. An A/B test splits live traffic between two versions. A hallucination is a fluent answer that is not supported by the facts the model was given.
The recurring test: measure before you change, and change the layer that failed. The official sample question has the pattern: confident wrong answers right after a document refresh, with the model and latency unchanged, point at retrieval, not at the model. Options that blame the model first, tune a setting unrelated to the symptom, or ship a change without a measurement are usually the distractors.
What the exam guide tests
The official CCAR-P guide lists 6 objectives for this domain. Exam questions are written against them, and so are the lessons: each row says where it is covered.
| # | Objective | Lesson |
|---|---|---|
| 4.1 | Define evaluation metrics (accuracy, latency, cost, safety, security) | 4.1 |
| 4.2 | Design evaluation datasets and test frameworks using mixed methodologies | 4.2 |
| 4.3 | Conduct A/B testing and iterative improvements | 4.3 |
| 4.4 | Diagnose system issues (prompt failure, hallucinations, model mismatch) | 4.4 |
| 4.5 | Optimize token usage, latency, and cost-performance trade-offs | 4.5 |
| 4.6 | Monitor system performance using logging and observability tools | 4.6 |
Lessons
- 4.1 Evaluation metrics: defining accuracy, latency, cost, safety and securityHow to write success criteria before a pilot, pick accuracy metrics by task type, measure latency, cost, safety and security, and choose release gates.22 min
- 4.2 Evaluation datasets and test frameworks: mixed methods that hold upHow to build a stratified eval set from real, edge, adversarial and synthetic cases, grade each metric by code, judge or human, and gate every change.21 min
- 4.3 A/B testing and iterative improvement: proving a change on live trafficHow to improve a Claude feature one tested change at a time: error analysis, offline evals, A/B test design, shadow and canary releases, and reading results.21 min
- 4.4 Diagnosing system issues: prompt failures, hallucinations and model mismatchTrace a bad answer to the layer that caused it: reproduce, isolate, compare, then tell prompt failures, hallucinations, lookalikes and model mismatch apart.23 min
- 4.5 Optimising tokens, latency and cost without losing qualityWhere Claude's tokens and money go, how to split batch from interactive work, which levers cut cost or waiting, and how to prove each saving kept quality.23 min
- 4.6 Monitoring a live Claude system: logs, dashboards, alerts and incidentsWhat each Claude call should log, which dashboards and alerts catch a quiet regression after a prompt change, and how to go from an alert to the failing step.23 min
Practice question
From the CCAR-P bank, tagged to this domain. Every answer option is explained. Nothing is stored, nothing to sign up for.
30 CCAR-P questions on this domain, free
The domain quiz in the question bank draws 10 random questions from the 30 tagged to Evaluation, Testing & Optimization, scores them and explains every option. Repeat it until the weak spots are gone, then sit the 63-question timed simulator.
Open the CCAR-P question bank → Start lesson 4.1 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.