Home › Study guides › CCAR-F › Domain 4 › Lesson 4.1
CCAR-F · Domain 4 · 20% of the exam · Lesson 4.1 · 20 min read
Explicit criteria: precision and false positives
Why "be conservative" does not cut false positives in an automated code review, how one noisy category costs trust, and how explicit criteria fix it.
Written against task statement 4.1 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
4.1.1 Why a reviewer that flags everything gets ignored
Picture hiring an inspector to look over a house before you buy it, with a one-line brief: "check that everything is in good shape". The report comes back with ninety items: a scuffed skirting board, a sticky drawer, a hallway painted "a slightly different white". Somewhere around item sixty, one line says a beam in the basement is cracked. The inspector did what you asked. "Good shape" never said what counted, so everything counted, and the one item that mattered is buried where you are least likely to read it.
Now the software version. A team runs Claude Code in its Continuous Integration (CI) pipeline, the automated jobs that run every time someone proposes a change to the code. That proposal is a pull request: the change waits there for review before it is merged into the main code. For every pull request, a job asks Claude Code to review the change, and a script posts each finding as a comment on the line it concerns. The review covers three categories: security, bugs and comment accuracy. Comments are the notes programmers leave in code to explain it, and the instruction for that category reads "check that comments are accurate".
Within a month the pull requests are full of comment-accuracy findings. A comment is "too terse", a TODO note (a reminder to finish something later) "may be outdated", a phrase is "informal". Most are not problems at all.
Developers start scrolling past the bot. One afternoon they scroll past a correct warning about SQL injection: a database query built from what a user typed, which an attacker can hijack to read or change data. The change gets merged. Two terms name what went wrong. A false positive is a finding that is not really a problem. Precision is the share of findings that are real: three real problems among ten comments is 30% precision.
The lever that raises precision is not a stronger adjective, a sterner warning or a bigger model. It is explicit criteria: instructions that say which issues count as findings and which do not, written as tests the model can apply one case at a time.
4.1.2 What a vague instruction leaves undecided
Start with the question that matters: why did a sensible-sounding instruction produce mostly noise? Because "accurate" is not one test; it is a family of them. A comment can be judged complete or incomplete, current or stale, well phrased or clumsy, right or wrong about what the code does. The instruction does not say which of those you mean, so the model has to choose. A model trying to be thorough reads the word broadly and flags every comment that could be better. From where it sits, that is doing the job well.
An explicit criterion replaces the goal with a test. Here is the example to remember: flag a comment only when the behaviour it claims contradicts what the code actually does. Apply it to the comment # retries up to 3 times sitting above a function that makes a single attempt: the claim and the code disagree, so it is a finding. Apply it to # quick fix, tidy later: it claims no behaviour, so there is nothing to contradict, and it is skipped. Every comment gets the same yes-or-no question, and the answer no longer depends on the model's taste in comments.
Anthropic's prompting guidance frames the rule this way: treat Claude as a brilliant but new employee who lacks context on your team's norms, and explain precisely what you want. Its golden rule is a practical check. If a colleague with little context would be confused by your prompt, Claude will be too. Hand "check that comments are accurate" to two new reviewers and you would get two different lists.
A goal versus a test
"Check that comments are accurate"
"Flag only if the claim contradicts the code"
# retries up to 3 times over one tryflagged# quick fix, tidy laterskipped4.1.3 Why "be conservative" is the wrong fix
Here is the tempting fix, and the one the exam expects you to reject. The team sees too many findings and adds a line to the prompt: "Be conservative. Only report high-confidence findings." It feels as if it should work, because fewer, surer findings sounds like higher precision. Run it, and the list gets a little shorter, but the comment-accuracy noise is still there. Meanwhile a subtle but real bug that the review used to catch has gone missing.
The reason is that the false positives were never low-confidence guesses. The model is quite sure the comment is terse and the TODO is old, and it is right about both. What went wrong is that you and the model disagree about whether terse comments and old TODOs are findings at all. Confidence measures how sure the model is of its own judgement; precision failed because its definition of a finding is not yours. Back to the house inspector: tell him "only mention things you are certain about" and the scuffed skirting board stays in the report, because he is certain it is scuffed.
So a confidence bar filters on the wrong axis. It keeps the confident noise and can throw away real issues the model was less sure of. It is also invisible: "conservative" is a threshold the model sets for itself, you cannot see where it sits, and nothing pins it in place from one run to the next. A categorical criterion ("report database queries built from user input; skip naming preferences") draws the line where you can read it, review it and change it.
Other quick fixes fail the same way. A larger model applies the vague instruction more skilfully but still has to guess your team's definition, and writing IMPORTANT in capitals makes the instruction louder, not clearer. Only a criterion changes what the model is looking for.
Three quick fixes, one real one
4.1.4 How one noisy category costs you the accurate ones
If the security category is precise, why does a noisy comment category matter? Surely developers can ignore the noise and read the rest. In practice, few people weigh each bot comment on its own merits. People form a view of the reviewer as a whole, and after a week in which most of its comments were not worth reading, the sensible move is to stop reading them. The accurate security finding arrives in the same format, from the same bot, beside eight comment nitpicks, and it gets the same glance.
You have met this in a kitchen. A smoke alarm that goes off every time someone makes toast ends up with its battery removed. It was right about fires all along; it lost its job because of the toast. That is the effect the exam tests: a category with a high false-positive rate undermines developers' confidence in the accurate categories too, because trust is given to the reviewer, not to each category separately.
How trust erodes
To find the culprit, measure precision per category. Every finding ends up fixed or dismissed, and counting those outcomes by category shows where the noise lives. The figures below are illustrative, invented for this example rather than measured, but they show the typical shape:
| Category | Findings developers acted on | Action while the prompt is fixed |
|---|---|---|
| Security | 9 of 10 | Keep posting |
| Bugs | 8 of 10 | Keep posting |
| Comment accuracy | 2 of 10 | Disable, rewrite its criteria, re-test |
The remedy is to temporarily disable the high false-positive category. The job stops asking for comment-accuracy findings, while security and bug findings keep flowing. Meanwhile the team rewrites that category's criteria and tests the new version against past pull requests where the right answer is known. Why not leave it running during the fix? Every noisy comment keeps spending trust, and trust comes back more slowly than it goes. Why not switch off the whole review? The security and bug findings are the reason the review exists, and they were never the problem.
4.1.5 Say what to report and what to skip
With comment accuracy paused, the team rewrites the criteria for the whole review. A good set names both sides of the line. Report the categories where a miss costs something real: bugs that produce wrong behaviour, and security problems. Skip what may be true but is not worth a developer's time on a pull request: minor style and local patterns.
A local pattern is a convention that one part of the code follows on purpose. One example is an older part of the code whose functions all report failure by returning an error number, instead of the more common approach of stopping with an error message. Another is test code that deliberately ignores errors. A reviewer meeting it for the first time calls it inconsistent, but the team chose it and will not change it in this pull request. Without a skip rule, the reviewer flags it on every pull request that touches that code.
The skip list matters as much as the report list, because a reviewer asked to find problems leans towards finding some. Claude Code's best-practices guide warns that a reviewer prompted to find gaps will usually report some even when the work is sound, and advises flagging only what affects correctness. Anthropic's own Code Review service for pull requests makes the same choice by default: bugs that would break production, not formatting preferences.
Here is the review section of the job's prompt. The lines that matter are the Comments rule, which carries the contradiction test, and the last two Skip lines, which protect local patterns and harmless comments. The first Skip line leans on a linter, the automatic style checker most teams already run in CI.
## Report
- Security: SQL or shell commands built from user input, missing
authorisation checks, secrets in code, personal data written to logs
- Bugs: logic that returns a wrong result or crashes on a realistic input
- Comments: ONLY when the behaviour a comment claims contradicts the code
## Skip
- Formatting and style the linter already checks
- Naming preferences and "could be more readable" suggestions
- Patterns used consistently in the surrounding code
- TODOs, terse comments, comments that describe intent, not behaviour
Notice what is missing: no "be careful", no "high confidence only". This is also how comment accuracy comes back. Its rule is now the contradiction test, and it returns once a re-test on past pull requests shows its findings are mostly real.
4.1.6 Severity that means the same thing on every run
One inconsistency is left. The job labels each finding critical, high, medium or low, so developers know what must be fixed before merging. On Monday a query built by gluing user input into SQL is labelled critical; on Thursday an almost identical one is medium. A missing empty-list check is high in one file and low in the next. Developers learn that the labels mean nothing and stop using them to decide what to fix first, which is the trust problem again, this time in the labels.
The cause is the one you already know: "critical" and "low" are words, and the model decides what they mean each time. The fix is also familiar, with one addition. Give each level an explicit definition AND a concrete code example of it. Think of choosing paint: "dark blue" means something different to everyone, while a swatch means one thing. The definition says where the line falls. The example is the swatch, a real case on the right side of the line. The model then classifies a new finding by comparing it with the examples rather than by feel.
| Level | Definition | Code example to include |
|---|---|---|
| Critical | Exploitable, or can leak or destroy data | cursor.execute("SELECT * FROM users WHERE id = " + user_id) with user_id from the request |
| High | Wrong result on a normal path | for i in range(len(items) - 1): silently skips the last item |
| Medium | Wrong only on an edge case | first = cart.items[0] with no check for an empty cart |
| Low | Misleading, but behaviour is correct | # retries up to 3 times above a single attempt |
Memorise the pattern, not these particular rows: every level gets a definition and at least one concrete code example, and each team writes its own. Anthropic's prompting docs call examples one of the most reliable ways to steer Claude's output and say they improve consistency. A severity label is output like any other.
4.1.7 The exam traps
Every mistake here leaves the definition of a finding to the model, or leaves the noise in front of developers. The exam presents them as a review that is noisy or ignored, then asks for the best improvement.
- ✗ "Check that comments are accurate." ✓ "Flag a comment only when its claimed behaviour contradicts what the code actually does." A test the model applies per comment, not a goal it interprets.
- ✗ Adding "be conservative" or "only report high-confidence findings" to cut false positives. ✓ Specific categorical criteria for what to report and what to skip. Confidence filters on certainty, so confident false positives survive and real findings can vanish.
- ✗ Moving to a bigger model, or putting IMPORTANT in front of the vague instruction. ✓ Change what the instruction says. A better reader, or a louder instruction, still leaves your definition to be guessed.
- ✗ Keeping a noisy category running and telling developers to ignore it. ✓ Disable it temporarily while its prompt is fixed. Developers do not ignore one category; they ignore the reviewer.
- ✗ Switching off the whole review because one category is noisy. ✓ Disable only the high false-positive category. The accurate security and bug findings are the value.
- ✗ Severity levels that are named but not defined. ✓ A definition and a concrete code example for each level, so the same kind of issue gets the same label.
4.1.8 Put it together: tune a review until developers read it
You now have every piece, from the vague goal that turns into noise to the severity examples that keep labels stable. The quickest way to believe it is to measure precision yourself, then break the fix and watch consistency go.
Explicit criteria are the first lever in this domain, and the rest build on them. When a criterion is clear but the model still misjudges borderline cases, few-shot examples (4.2) show it how to handle them. Structured output (4.3) turns each finding into fields a script can post, and a detected_pattern field (4.4) shows which code constructs developers keep dismissing, which is how you spot the next noisy category. Confidence still has a job, just not this one: in multi-pass review (4.6), the model reports its confidence alongside each finding so that calibrated routing can decide what a human checks.
Key takeaways
- ✓ Precision is the share of findings that are real problems; a false positive is a finding that is not a real problem.
- ✓ Explicit criteria turn a goal into a test: "flag a comment only when its claimed behaviour contradicts the code", not "check that comments are accurate".
- ✓ "Be conservative" and "only report high-confidence findings" filter on certainty, not on what counts, so confident false positives survive and real findings can drop out.
- ✓ A high false-positive category undermines trust in the accurate ones, because developers judge the reviewer as a whole.
- ✓ Temporarily disable the noisy category, keep the precise ones running, and re-enable it after its criteria are rewritten and tested.
- ✓ Write criteria that say what to report (bugs, security) and what to skip (minor style, local patterns).
- ✓ Give every severity level explicit criteria and a concrete code example so classification is consistent.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
72 CCAR-F questions on Domain 4, free
Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 4 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.