Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 4

CCAR-P · Domain 4 of 7 · 6 lessons · about 133 min

Domain 4: Evaluation, Testing & Optimization

Measuring and improving a Claude solution: metrics, evaluation datasets, A/B tests, diagnosing failures, cost and latency, and monitoring. 16% of CCAR-P.

16%Of the exam
~10Questions on exam day
6Free lessons
30Practice questions here

This domain is 16% of CCAR-P, about 10 of the 63 questions. It covers how you know a Claude solution works, how you find out why it stopped working, and how you make it cheaper and faster without making it worse.

The vocabulary: an evaluation (eval) runs a fixed set of test cases through the system and scores the results. Code-based grading checks answers mechanically; LLM-as-judge has a model score answers against a rubric; human evaluation is the reference both are checked against. An A/B test splits live traffic between two versions. A hallucination is a fluent answer that is not supported by the facts the model was given.

The recurring test: measure before you change, and change the layer that failed. The official sample question has the pattern: confident wrong answers right after a document refresh, with the model and latency unchanged, point at retrieval, not at the model. Options that blame the model first, tune a setting unrelated to the symptom, or ship a change without a measurement are usually the distractors.

What the exam guide tests

The official CCAR-P guide lists 6 objectives for this domain. Exam questions are written against them, and so are the lessons: each row says where it is covered.

#ObjectiveLesson
4.1Define evaluation metrics (accuracy, latency, cost, safety, security)4.1
4.2Design evaluation datasets and test frameworks using mixed methodologies4.2
4.3Conduct A/B testing and iterative improvements4.3
4.4Diagnose system issues (prompt failure, hallucinations, model mismatch)4.4
4.5Optimize token usage, latency, and cost-performance trade-offs4.5
4.6Monitor system performance using logging and observability tools4.6

Lessons

  1. 4.1 Evaluation metrics: defining accuracy, latency, cost, safety and securityHow to write success criteria before a pilot, pick accuracy metrics by task type, measure latency, cost, safety and security, and choose release gates.22 min
  2. 4.2 Evaluation datasets and test frameworks: mixed methods that hold upHow to build a stratified eval set from real, edge, adversarial and synthetic cases, grade each metric by code, judge or human, and gate every change.21 min
  3. 4.3 A/B testing and iterative improvement: proving a change on live trafficHow to improve a Claude feature one tested change at a time: error analysis, offline evals, A/B test design, shadow and canary releases, and reading results.21 min
  4. 4.4 Diagnosing system issues: prompt failures, hallucinations and model mismatchTrace a bad answer to the layer that caused it: reproduce, isolate, compare, then tell prompt failures, hallucinations, lookalikes and model mismatch apart.23 min
  5. 4.5 Optimising tokens, latency and cost without losing qualityWhere Claude's tokens and money go, how to split batch from interactive work, which levers cut cost or waiting, and how to prove each saving kept quality.23 min
  6. 4.6 Monitoring a live Claude system: logs, dashboards, alerts and incidentsWhat each Claude call should log, which dashboards and alerts catch a quiet regression after a prompt change, and how to go from an alert to the failing step.23 min

Practice question

From the CCAR-P bank, tagged to this domain. Every answer option is explained. Nothing is stored, nothing to sign up for.

30 CCAR-P questions on this domain, free

The domain quiz in the question bank draws 10 random questions from the 30 tagged to Evaluation, Testing & Optimization, scores them and explains every option. Repeat it until the weak spots are gone, then sit the 63-question timed simulator.

Open the CCAR-P question bank → Start lesson 4.1 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.