Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 4 › Lesson 4.6

CCAR-F · Domain 4 · 20% of the exam · Lesson 4.6 · 21 min read

Multi-instance and multi-pass review

Why the session that wrote code is a weak reviewer of it, how per-file and integration passes fix large reviews, and how confidence routes findings.

Written against task statement 4.6 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

4.6.1 Why a good review needs someone who wasn't there

Most workplaces that care about quality run on one quiet rule: whoever did the work doesn't sign it off. An accountant's figures go to an auditor. A contract drafted by one lawyer is read by another before anyone signs it. Nobody assumes the first person is careless. The rule exists because the person who did the work reads it through what they meant to do, and the gap between what they meant and what they produced is exactly what they cannot see.

Now put Claude Code into a continuous integration (CI) pipeline: the automated checks that run on every pull request, a proposed code change, before it joins the main code. Two setups look efficient and quietly fail. In the first, Claude Code writes a change and the same session is then asked to review it; it approves code that contains a bug. In the second, one prompt reviews an 11-file pull request in a single pass; it goes deep on some files, skims others and contradicts itself about identical code. Neither problem goes away with a better-worded prompt, because neither is a wording problem.

Both are problems of context: what the reviewer is carrying when it looks at the code. So the fix is architectural. A multi-instance review hands the code to a separate Claude instance that never saw the reasoning behind it. A multi-pass review splits one oversized review into several focused ones, each with a narrow context and a single question. A final verification pass has the model attach a confidence to every finding, so people with limited review time look at the right ones first.

4.6.2 Why the session that wrote the code approves its own bug

Start with the first failure. A pipeline hands Claude Code a ticket for the restocking service: send a reorder alert when a product has 10 or fewer units left. While planning, the session restates the ticket to itself as "alert when stock drops below 10" and writes the check if stock < 10 ("if stock is less than 10"). With exactly 10 units on the shelf, no alert goes out. The pipeline then asks the same session to review its change for bugs. The answer comes back: the logic matches the requirement, approved.

Nothing about that review was lazy. The model's context is everything it can see while it works, and in this session that means the ticket, the plan, the reasoning, the assumptions and the code. When the model reviews, it reads the code alongside all of that, and its own restatement, "below 10", sits right there looking like the requirement. So it checks the code against its plan, and the code matches the plan perfectly. The bug lives in the plan, the one thing the author has no reason to question.

That is the self-review limitation: a model keeps the reasoning context from generation, which makes it less likely to question its own decisions in the same session. You know it from writing directions to your home for a friend. Read back, they look flawless, because your knowledge fills every vague step: "left after the bakery" plainly means the old bakery on the corner. A stranger holding only the page takes the wrong left. The author's context fills the gaps in the work, so the author cannot see them; a reviewer without that context meets the work as it actually is.

What each reviewer has in context

Same session self-review

The ticket"10 or fewer units"
Its own plan"alert below 10"
Its reasoning and assumptions
The code it wrotestock < 10

checks the code against its plan

Independent instance fresh context

The ticket"10 or fewer units"
The diffstock < 10
The review criteria
No plan, no reasoning

checks the code against the ticket

The authoring session reviews the code through its own plan; an independent instance meets the code and the ticket with nothing in between.

4.6.3 Fresh eyes: the independent review instance

Here is the question that trips people up: if self-review misses the bug, what is the reliable fix? Two answers are tempting because they keep everything in one session, and both fail for the same reason. The first is a sterner instruction: "review your changes critically, as if a stranger wrote them." An instruction changes what you ask, not what the model is holding. The plan that says "below 10" is still in the context, and the model still reads the code through it.

Asking Claude to check its work is not useless; it does catch slips such as a typo or a forgotten case. What it cannot do is make the model unsee its own reasoning, so the subtle errors that sit inside that reasoning survive.

The second tempting answer is extended thinking, the setting that lets Claude reason step by step before it answers. More reasoning sounds like more scrutiny, but it starts from the same context and the same premise. Thinking harder about "alert below 10" produces a longer, more confident case that stock < 10 is correct.

The fix is an independent review instance: a second Claude instance whose context holds the diff (the lines the change added and removed), the ticket and the review criteria, and none of the generator's reasoning. The exam guide calls the principle session context isolation. Independence is a property of the context, not of the model; the reviewer can be the very same model.

In CI the reviewer can be a new run of claude -p, Claude Code's non-interactive mode for scripts and pipelines. It can also be a fresh subagent (a helper that Claude Code starts with an empty context of its own) or a separate API call. It cannot be a --continue of the generating run, which reopens that run's conversation, or a forked subagent, which starts with a copy of it. And whatever you hand the reviewer, leave out the author's explanation of the change: pasting it in imports the bias you were leaving behind.

Proposed fix What the reviewer has in context Likely to catch stock < 10?
Same session, told to "review critically" The ticket, its plan ("below 10"), its reasoning, the code Unlikely: it checks the code against its own plan
Same session, extended thinking on All of that, plus longer reasoning built on the plan Unlikely: more thought from the same premise
Independent instance The diff, the ticket, the review criteria Likely: it compares < 10 directly with "10 or fewer"

Memorise the last row and the reason behind the first two: both keep the reviewer inside the author's context.

Here is one hypothetical way to wire it as two steps of a CI job. The line to watch is the pipe (|) that feeds the diff into a brand-new claude -p run. The commented-out --continue line is the mistake, because it reopens the author's conversation.

# Step 1: Claude Code writes the change. Its plan and reasoning stay in this run.
claude -p "Implement ticket STK-212: reorder alert at 10 or fewer units" \
  --allowedTools "Read,Edit"

# Step 2: a NEW run reviews it. It sees the diff, the ticket and the criteria only.
git diff main | claude -p \
  "Review this diff against ticket STK-212 (alert at 10 or fewer units). Report bugs only." \
  --output-format json --json-schema "$FINDINGS_SCHEMA"   # findings as structured JSON

# Not this: --continue reopens Step 1's conversation, so the author reviews itself.
# claude -p "Now review your change for bugs" --continue

4.6.4 One pass over many files: attention dilution

The second failure appears when the change is big. A pull request changes 11 files across the restocking service, and the pipeline reviews them all together in one prompt. The review that comes back is lopsided. The first few files get detailed, line-by-line comments; the later ones get a sentence each, and an obvious bug in one of them slips through. Worse, the review contradicts itself. A division by the day's sales count, which crashes on a day with no sales, is flagged in sales_velocity.py and waved through, character for character the same, in restock_forecast.py.

The cause is attention dilution. All 11 files fit comfortably in the context window, the most text the model can take in at once, so space is not the problem; attention is. When one call must judge thousands of lines at once, each file gets only a share of the model's focus, and the shares are uneven. Claude Code's documentation states the general rule: performance degrades as the context fills. Contradictions follow naturally. The model reaches its verdict on the ninth file with different focus, and different material around it, than its verdict on the second, so nothing holds the two to the same standard.

Picture a food inspector asked to check a dozen restaurant kitchens in one afternoon. The first kitchen gets thermometers in the fridges and a torch under the sinks. By the tenth it is a glance from the doorway, and a cracked tile that earned a warning in kitchen 3 passes unremarked in kitchen 11. The inspector didn't get worse at inspecting; the job was cut too thin. In model terms: one call spread over many files reviews each of them less deeply and less evenly.

That is also why telling the single pass to "give every file equal care" doesn't fix it: the instruction leaves the job exactly as big. The fix is multi-pass review, which changes the shape of the job instead of asking for more effort or a bigger model:

  1. PER-FILE PASSES. One focused call per file, each with that file's changes and the same review criteria, asking only about local issues: logic, edge cases and error handling inside that file. Every file gets full attention, and identical code meets identical criteria, so the division by zero is flagged in both places.
  2. INTEGRATION PASS. One separate call that sees the per-file findings plus a short note of what each file sends to and expects from the others. It asks a single question: does data still flow correctly between these files? This is where the review notices that warehouse_sync.py now reports stock in cases of 12 while reorder_alert.py still reads the number as single units. Neither file is wrong on its own.
  3. MERGE. Your pipeline combines both sets of findings and removes duplicates, so each issue is reported once.

A multi-pass review of the 11-file pull request

The pull request11 changed files
11 per-file passesone file, same criteria
Integration passdata flow between files
Merged findingsdeduplicated, one per issue
Eleven narrow passes find local issues with full attention, one integration pass asks what no single file can show, and your pipeline merges the results.

In CI this is a loop that calls claude -p once per file, then once more for the integration pass. The per-file passes don't depend on each other, so they can run in parallel. The integration pass has to wait, because it needs their findings.

4.6.5 Verification passes: confidence that routes attention

Multi-pass review solves one problem and creates a pleasant new one: more findings. Eleven per-file reports plus an integration report can hold dozens of candidates, and not all of them are real. Post everything and developers drown in noise; worse, a stream of false alarms teaches them to ignore the true ones too. Your human reviewers have perhaps an hour a day. The question becomes which findings deserve that hour.

The answer is a verification pass: one more pass that takes each candidate finding, checks it against the actual code, and reports a confidence alongside it. The model self-reports how sure it is that each finding is real, as a field in its structured output, and your pipeline uses that number to decide where the finding goes. Anthropic's managed Code Review service works along similar lines: several agents look for different kinds of issue in parallel, then a verification step checks each candidate against actual code behaviour before anything is posted.

In the running example, the boundary finding comes back with high confidence, because the ticket and the code disagree in black and white. A vaguer finding, "two orders might reserve the last unit at the same moment", comes back low, because proving it needs knowledge of the database that the pass cannot see. Your pipeline posts the first as a comment and sends the second to a person to check.

One word in this pattern carries the weight: calibrated. A confidence is calibrated when it matches reality: of the findings marked 90% sure, about 90% turn out to be real. Raw self-reported confidence is poorly calibrated; a model can be sure and wrong. So you don't trust the number as it stands. You run the verification pass over past findings whose outcome you already know, real bug or false alarm, and see how often findings at each confidence level turned out real. Those results set the thresholds, and only then does "high confidence" mean something.

Calibrated confidence Where the finding goes Why
High Posted automatically as an inline comment on the pull request Past findings at this level were nearly always real
Medium A human reviewer checks it before it is posted Real often enough to matter, wrong often enough to need a person
Low A triage list that a person skims; never deleted Mostly false alarms in the past, but some were real bugs

The bands are your design choice; what matters is that the thresholds come from labelled past findings. Confidence decides WHO looks at a finding, not WHETHER it exists. That is what separates it from the instruction "only report findings you're highly confident about". That instruction silently discards findings and does not make the model more precise; specific review criteria do that job.

4.6.6 The exam traps

Most traps in this task statement offer more effort in the same place instead of a different shape of review. Find the option that leaves the reviewer's context unchanged, and you have usually found a distractor.

  • ✗ Adding "review your changes critically" to the session that wrote the code. ✓ Review in an independent instance with the diff, the requirements and the criteria. The instruction changes the request but not the context, so the author's plan still frames the review.
  • ✗ Turning on extended thinking so the author reviews more deeply. ✓ Use a fresh instance. More reasoning inside the same context builds on the same premise instead of questioning it.
  • ✗ A model with a bigger context window, so one pass covers every file. ✓ Per-file passes plus an integration pass. The files already fit; the failure is attention spread too thin, not a lack of space.
  • ✗ Three independent full reviews of the pull request, keeping only issues at least two of them report. ✓ Focused passes. Voting suppresses real bugs that are caught only some of the time, and every run still suffers the same dilution. The word "independent" in this option is a lure: independence helps only when it removes the author's context.
  • ✗ Asking each per-file pass to "also check interactions with other files". ✓ A separate integration pass. A pass that sees one file cannot see the other end of a data flow.
  • ✗ Filtering with "only report high-confidence findings", or trusting raw scores. ✓ Have a verification pass attach a confidence to every finding. Calibrate the thresholds on labelled past findings and route the uncertain ones to a person, so nothing is dropped on a guess.

Three tempting fixes, one that works

Bigger context windowthe files already fit
Three runs, keep consensusvotes out bugs caught once
Make developers split PRsmoves the work, fixes nothing
Per-file passes plus an integration passfull attention per file, one pass for data flow
The single-pass failure is attention spread too thin, so only changing the shape of the review fixes it.

4.6.7 Put it together: build a review pipeline and break it

You now have every piece: the self-review limitation and its fix, attention dilution and its fix, and a verification pass that routes what remains. The fastest way to believe all of it is to watch each part fail.

The confidence field you added in step 3 is only the start of calibration, and Domain 5 finishes the job. Human review workflows (5.5) build labelled validation sets, keep sampling high-confidence output and check accuracy segment by segment before any finding skips a person. And the reason per-file passes work, keeping each context small and focused, runs through the rest of Domain 5, from long conversations (5.1) to exploring large codebases (5.4).

Key takeaways

  • ✓ The session that generated code keeps its plan and reasoning in context, so it checks the code against its own intentions and rarely questions them.
  • ✓ An independent instance with only the diff, the requirements and the criteria catches subtle issues that self-review instructions and extended thinking miss.
  • ✓ Independence is about context: a new claude -p run, a fresh subagent and a separate API call all qualify; --continue, a fork and the author's notes do not.
  • ✓ One pass over many files dilutes attention, producing uneven depth, missed bugs and contradictory verdicts on identical code, even when every file fits in the context window.
  • ✓ Split large reviews into per-file passes for local issues plus a separate integration pass for cross-file data flow, then merge the findings.
  • ✓ A verification pass has the model report a confidence with each finding, and thresholds calibrated on labelled past findings route each one to posting or to a person, never to deletion.
  • ✓ Bigger windows, consensus voting, sterner instructions and "only report high-confidence findings" are the distractors; the fix changes the shape of the review, not its effort.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

72 CCAR-F questions on Domain 4, free

Every question in the bank is tagged to a domain, so you can drill 72 questions on Prompt Engineering & Structured Output alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 4 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources