Home › Study guides › CCAR-F › Domain 3 › Lesson 3.5
CCAR-F · Domain 3 · 20% of the exam · Lesson 3.5 · 23 min read
Iterative refinement: examples, tests and the interview pattern
How to steer Claude Code from a first attempt to a right one with input/output examples, tests written first, the interview pattern and well-batched fixes.
Written against task statement 3.5 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.5.1 Why the first answer is rarely the last one
Picture a developer on a team that has adopted Claude Code. The company is moving its customers from an old system to a new one. The old system exports each customer as a record (a name, a phone number, a signup date and so on), and the new system expects those records in a different shape. Someone has to write the function, a small named piece of code, that takes one old record and returns one new record. The developer types "write a function that converts customer records from the old export format to the new one" and gets back tidy, plausible code.
It is also wrong. It splits names at the last space, so "Maria de la Cruz" becomes first name "Maria de la" and last name "Cruz". It drops the country code from phone numbers. Where the old export had no value, it writes the word "None" as text instead of leaving the field empty. A second attempt, with a longer description, is wrong in different ways.
Claude Code is not malfunctioning here. A description in English has more than one reasonable reading, and the model picks one of them each time. It cannot see the records in your head, and it cannot know which of the twenty small decisions hidden in "convert" you care about. Anthropic's prompting guide suggests thinking of Claude as a brilliant but new employee who lacks context on your norms. You would not hand that colleague a one-line brief and be surprised when the result needs another round.
So the skill here is not "write the perfect prompt". It is iterative refinement: getting from a first attempt to a right one in as few, well-aimed rounds as possible. There are four techniques, each matched to a situation. When prose keeps being misread, you give concrete input/output examples. When you want steady, measurable improvement, you write the tests first and share the failures. When the territory is unfamiliar, you use the interview pattern and let Claude ask the questions. And when a result has several problems, you decide whether to send them in one message or one at a time, depending on whether the fixes interact. We will follow the customer record converter through all four.
3.5.2 Show, don't describe: two or three examples
The first failure is the one you just saw: the description sounds fine, but every attempt reads it differently. Each converter is a defensible reading of the sentence. Claude Code did not ignore you; the sentence under-specifies the job, and every rewrite of it leaves a different gap.
The fix is to stop describing the transformation and start showing it. Give Claude Code two or three concrete pairs: this exact old record in, this exact new record out. An example carries many decisions at once, and it carries them without ambiguity.
In the pairs below, the first shows that a phone number keeps its country code and loses its spaces. In the second, "Maria de la Cruz" comes out as first name "Maria" and last name "de la Cruz", so the rule "split at the first space" is settled without anyone having to phrase it. The third settles the empty-value rule the same way: an empty phone becomes null, the standard marker for "no value here".
IN: {"name": "Tom Baker", "phone": "+44 20 7946 0018"}
OUT: {"first_name": "Tom", "last_name": "Baker", "phone": "+442079460018"}
IN: {"name": "Maria de la Cruz", "phone": "+1 415 555 0100"}
OUT: {"first_name": "Maria", "last_name": "de la Cruz", "phone": "+14155550100"}
IN: {"name": "Sam Okafor", "phone": ""}
OUT: {"first_name": "Sam", "last_name": "Okafor", "phone": null}
Think of explaining a knitting pattern over the phone versus handing someone the finished sleeve. The sleeve answers questions the caller did not know to ask. Precisely: an input/output example is the most effective way to communicate a transformation when a prose description is interpreted inconsistently, because it specifies the outcome instead of paraphrasing it. Choose examples that differ from each other, as these three do: an ordinary record, a surname of several words, a missing phone. Three near-identical examples teach one rule three times; three varied examples teach three rules.
A prose description versus two or three examples
Prose description
Two or three examples
the transformation is now unambiguous
Anthropic's prompting guide adds three tips that apply here too. Make the examples relevant to the real use case, and diverse enough to cover the unusual cases. Mark them clearly, for example inside <example> tags, so Claude can tell an example from an instruction. In Claude Code you can paste them straight into your message, or keep them in a file and reference it with @, which pulls the file's full content into the conversation.
3.5.3 Write the tests first, then share what fails
Examples fix a misread specification. The next problem is different: the converter is roughly right, and you want it to get steadily better without re-reading every line after every change. Judging code by eye does not scale. The Claude Code docs are blunt about this: without a check Claude can run, "looks done" is the only signal available, and you become the verification loop.
The technique for this is test-driven iteration. A test is a few lines of code that feed the converter one input and compare what comes out with the output you expect. It passes or it fails, and most test runners (the programs that run tests) colour passes green and failures red. A test suite is the whole collection of tests for one piece of code, run together with a single command. Test-driven iteration means writing that suite BEFORE asking for any implementation, and the suite covers three things:
- Expected behaviour. An ordinary record converts to exactly this output.
- Edge cases. Unusual but legitimate inputs at the edge of the normal range: an empty phone, a name with no space, a missing date. First versions break here most often, because they were written with the ordinary case in mind.
- Performance. How fast it must be: say, 100,000 records converted in under ten seconds.
At first every test fails, because there is no code yet. That is the point. You have turned "a good converter" from an opinion into a list of red and green lights.
Then you ask Claude Code to make the tests pass. Claude Code can run the suite itself when you ask it to and read the results, or you can run it and paste the output. Either way, what drives the next round is the actual failure, not "it's still wrong": which test failed, what it expected, what it got. A failing test is the most specific instruction you can give, because it names the exact input and the exact expected output. Claude Code changes the code, the suite runs again, and the loop repeats until nothing is red. Each round is a measured step forward, which is what progressive improvement means here.
Test-driven iteration
One habit keeps this honest. Tell Claude Code that the tests are there to verify correctness, not to define the solution. Otherwise it may write code that recognises the exact test records and returns the expected answers for them, which passes the suite and still fails on real data. Anthropic's prompting guide suggests exactly that wording. The other half of the habit comes from the Claude Code docs: ask for evidence, meaning the test output itself, rather than a claim that "the tests pass".
3.5.4 The interview pattern: let Claude ask first
Examples and tests both assume you know what right looks like. Now suppose the migration will take months, and until it finishes the new system must fetch customer details from the old one, which is slow. The team wants a cache: a stored copy of recent answers, so the new system does not have to ask the old one every time. Nobody on the team has built a cache before.
You cannot write a good test suite for something you do not understand, and you cannot give examples of behaviour you have not thought about. The danger here is not a misread prompt. It is a prompt that leaves out considerations you did not know existed, so the implementation quietly makes those decisions for you.
The interview pattern turns the conversation around. Instead of describing the cache and asking for it, you ask Claude Code to interview you before it implements anything. It asks its questions with AskUserQuestion, a built-in Claude Code tool that shows you a multiple-choice question; you pick an option or type your own answer. The Claude Code docs suggest a prompt of this shape. The parts that matter are the request for an interview, the push toward the hard parts, and the spec (short for specification, the written list of decisions) at the end.
I want to add a cache in front of the old customer system. Interview me in
detail using the AskUserQuestion tool. Don't ask obvious questions; dig into
the hard parts I might not have considered. Keep going until we've covered
everything, then write a complete spec to SPEC.md.
The questions that come back are the value. When the old system changes a customer's phone number, how does the stored copy find out? That is cache invalidation: deciding when a stored copy is out of date and must be thrown away. What happens when the old system is down: serve a possibly stale copy, or show an error? That is a failure mode, one of the ways the system can go wrong, together with what it should do when it does. How big can the cache grow, and what gets removed first when it is full?
It is the difference between telling a builder "put in a kitchen" and having the builder walk the room with you, asking where the water comes in. You do not need to know the plumbing; you need someone who does to ask you the right questions. Precisely: the interview pattern has Claude surface design considerations, such as cache invalidation strategies and failure modes, before implementation. It is the technique for unfamiliar domains, where the developer cannot anticipate those considerations alone.
The interview pattern
AskUserQuestionTwo details from the docs are worth keeping. First, once the spec is written, start the implementation in a fresh session. A session is one continuous conversation with Claude Code; a fresh one starts empty, so Claude works from the clean spec rather than from the long interview that produced it. Second, the docs recommend the interview for larger features. For a small change whose right answer you already know, it is overhead.
3.5.5 One specific test case beats a paragraph of complaint
Back to the converter, which now sits inside a migration script: a program that runs once over every record in the old export and loads the results into the new system. It works on the sample data and then fails on the real thing, because real data has holes. Where the old export had no phone number, the new records contain the text "None"; where a signup date was missing, the script crashes.
The developer's first instinct is to write "handle nulls properly" (a null is an empty value, a field with nothing in it). That is the prose problem all over again. Properly how? Skip the record, store a null, store empty text, or invent a default date?
The better message is a specific test case: one example input that contains the missing values, and the exact output you expect. Here the examples technique and the tests technique meet in a single message. The test is written in Python, where None without quotes is the language's null and "None" in quotes is just a four-letter word. Each assert line states one thing that must be true, and the test fails if it is not; the first two are the decision.
# test_convert_customer_record.py (paste this to Claude Code, not a description)
def test_missing_phone_and_signup_date():
old = {"id": 4471, "name": "Maria de la Cruz", "phone": "", "signed_up": None}
new = convert_customer_record(old)
assert new["phone"] is None # empty phone -> null, never the text "None"
assert new["signed_up"] is None # missing date -> null, never a crash or a default
assert new["first_name"] == "Maria" # the rest of the record still converts
Notice what this does that "handle nulls" cannot. It removes the ambiguity between empty text, a real null and the word "None". It declares that a missing date is not an error. And it stays in the suite as a permanent guard, so the bug cannot quietly return. One such test per edge case, added as you find them, is how a migration script goes from "works on the sample" to "works on the export".
3.5.6 One message or one at a time?
The last decision comes after review. The converter now has three known problems. It reads signup dates in the wrong format, taking 03/04/2021 as the 3rd of April when the old system meant March 4th. It applies the wrong time zone to them. And the lines it writes to its log (the running record of what the script did) start with inconsistent labels. Do you send all three in one message, or fix them one by one?
The rule is whether the fixes interact, meaning fixing one changes what the other has to do. The date-format fix and the time-zone fix both rewrite the same few lines that turn the text of a date into a real date. Fix the format alone and the time-zone logic now works on different values; fix the time zone alone and the later format change can undo it. Fixes that interact go in one detailed message, so Claude Code sees both constraints at once and solves them together.
The log label problem touches none of that. It is independent, and independent problems are best fixed one round at a time. Small rounds are easier to check, easier to undo if one goes wrong, and each change can be reviewed on its own.
Think of a mechanic. If the brakes pull to one side and the wheel alignment is off, you describe both in one visit, because adjusting one changes the other. A squeaky door is a separate job; mention it when the brakes are done.
| Problem | What fixing it changes | How to send it |
|---|---|---|
| Dates read in the wrong format | The date values the time-zone step works on | With the time-zone fix, in one detailed message |
| Wrong time zone applied | The same date values, once they are read | With the format fix, in that same message |
| Inconsistent log labels | Only the log lines | On its own, in a separate round |
Precisely: ask one question of every pair of problems. Does fixing one change what the other must do? Yes means one detailed message; no means separate rounds, one issue each.
3.5.7 The exam traps
Every technique in this lesson has a tempting wrong twin, usually "explain it again, harder".
- ✗ Rewriting the description a fourth time when results keep varying. ✓ Give 2 to 3 varied input/output examples. The variation comes from the description having several readings; examples remove the readings, a reword just changes them.
- ✗ Implementing first and asking for tests afterwards. ✓ Write the tests first, covering expected behaviour, edge cases and performance, then iterate on failures. Tests written after the fact describe what the code does, not what it should do.
- ✗ Telling Claude Code "it's still wrong" or "handle nulls properly". ✓ Share the concrete failing test output, or one specific test case with input and expected output. A precise failure is an instruction; a complaint is not.
- ✗ Writing a detailed specification up front for a domain you do not know. ✓ Use the interview pattern. You cannot specify considerations you have not anticipated; Claude's questions surface them.
- ✗ Sending every issue in one message regardless of how they relate. ✓ One message only when the fixes interact; otherwise sequential rounds. Bundling unrelated fixes makes the change hard to review.
- ✗ Fixing interacting issues one at a time and watching each fix undo the last. ✓ Put them together in one detailed message. If fixing A changes B, the model needs both constraints at once.
Put side by side, the techniques form one map from symptom to fix:
| Situation you notice | What to send | Why it works |
|---|---|---|
| The description is read differently every time | 2 to 3 concrete input/output examples | Examples specify the outcome, not a paraphrase of it |
| You want steady, measurable improvement | A test suite written first, then the failures | Each failure is a precise instruction with input and expected output |
| The domain is unfamiliar to you | An interview before any implementation | Claude's questions surface considerations you could not list |
| One edge case is mishandled (nulls in a migration) | One specific test case: example input, expected output | Names the case and the decision in a form Claude Code can run |
| Several problems, and the fixes affect each other | One detailed message with all of them | Claude Code can solve them together instead of undoing itself |
| Several problems, unrelated to each other | One issue per round, in sequence | Each round stays small, checkable and easy to undo |
Memorise the mapping from the first column to the second; the exam presents the situation and asks for the technique. The third column is the reasoning that lets you recognise a reworded version.
3.5.8 Put it together: refine the converter until it is right
You now have every technique and the map between them. The fastest way to make the mapping stick is to run the converter through each technique and, at the end, watch what happens when you use the wrong one.
Everything here assumes a developer at the keyboard who can answer questions and paste failures. Claude Code in CI/CD (3.6) removes that person. A pipeline run has nobody to interview and nobody to correct it, so the examples, tests and expected outputs you supplied by hand must be in the prompt and the pipeline from the start. Few-shot prompting (4.2) then takes the examples idea into the prompts of your own applications, where a few worked examples keep the output consistent on ambiguous cases.
Key takeaways
- ✓ Iterative refinement means each round hands Claude Code something more concrete than the last, not the same description reworded.
- ✓ When a prose description is interpreted inconsistently, 2 to 3 varied input/output examples are the most effective way to specify the transformation.
- ✓ Test-driven iteration writes the test suite first, covering expected behaviour, edge cases and performance, then improves the code by sharing the concrete failures.
- ✓ The interview pattern has Claude ask questions before implementing, surfacing considerations such as cache invalidation and failure modes in an unfamiliar domain.
- ✓ A mishandled edge case, such as nulls in a migration script, is fixed with one specific test case: example input and expected output.
- ✓ Interacting issues go in one detailed message so the fixes are solved together; independent issues are fixed one at a time.
- ✓ Wrong answers keep describing, or apply a real technique to the wrong symptom; the right answer hands Claude something concrete that fits the symptom.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
72 CCAR-F questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 72 questions on Claude Code Configuration & Workflows alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.