Home › Study guides › CCAR-F › Domain 5 › Lesson 5.4
CCAR-F · Domain 5 · 15% of the exam · Lesson 5.4 · 24 min read
Context in large codebase exploration
Why a day-long codebase exploration drifts into vague answers, and the fixes: scratchpad files, subagents, phase summaries, manifests and /compact.
Written against task statement 5.4 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
5.4.1 Why a long exploration ends in vague answers
Picture your first week in a job that runs on paperwork. Your manager hands you a room of old supplier contracts and asks questions as they come up. At nine you find the late-delivery penalty in contract 17, clause 4, and answer precisely. By five you have skimmed two hundred more contracts, and when the question comes back you say "contracts like these usually put penalties near the end". The fact isn't gone; it's buried under everything you read after it, so you reach for the general shape instead.
An agent exploring a large codebase, all the code behind one system, hits the same wall. Everything it reads lands in its context window, the text the model sees on every turn. Exploring is mostly reading, so the window fills with discovery output: pages of search hits, whole files opened to check one function, long directory listings. The good findings sit in the same window as all of that, and every hour there is more of it.
The answer is not to hope the model remembers harder; it is a plan for where knowledge lives. Key findings go into files the agent re-reads. Heavy reading goes to helper agents called subagents, each with its own context window, so only their answers come back. Short summaries carry what one stage of the work learned into the next, and progress is saved as the work goes. When the window does fill with old search output, a command called /compact shrinks it by summarising.
Only two of these are built into Anthropic's tools: subagents and the /compact command. The files of findings, the summaries and the saved progress are design patterns. They are ordinary files and prompts that you plan and tell the agent to use; no setting switches them on.
5.4.2 The symptom: typical patterns instead of specific classes
Here is the failure this task statement is built around. Our running example: you are building a developer productivity tool with the Claude Agent SDK, Anthropic's software development kit (SDK), a ready-made library for building agents. The tool is an agent that helps engineers understand unfamiliar systems. It uses the SDK's built-in tools: Read opens a file, Write saves one, Grep searches inside files for text, Glob finds files by name and Bash runs commands. An engineer points it at a hypothetical legacy order-fulfilment system (fifteen years old, a few thousand files, no documentation) and asks it questions all day.
In the morning the answers are sharp. They name specific classes, the named building blocks the code is organised into, and the file each one lives in. Refunds are posted by RefundLedgerService.post_reversal() in billing/rl_svc.py; stock is reserved by WarehouseAllocator in stock/alloc_v2.py. By mid-afternoon the agent has run hundreds of searches and read eighty files. Asked what a refund does to stock, it replies "Typically, a refund service calls an inventory component to put the item back". Asked again where stock is reserved, it says "probably in the order controller", contradicting its own morning answer.
This is context degradation, and it has two tell-tale signs. The agent describes typical patterns, how systems like this usually work, instead of the specific classes it found. And it gives inconsistent answers, contradicting what the same session established hours earlier. Both mean the specifics are no longer reliably within reach, so the model fills the gap with general knowledge of order systems: fluent, plausible and wrong for this codebase.
Two forces push the specifics out of reach. As the context grows, the model's ability to recall a particular detail from it declines gradually; there is no cliff edge, just less precision across a long, crowded input. And near the limit, the Agent SDK compacts the conversation automatically: it replaces older history with a summary to free space. A summary keeps the gist, but nothing guarantees it keeps every class name and file path from nine o'clock.
The same session at 9am and at 4pm
9am context
RefundLedgerService, billing/rl_svc.py4pm context
5.4.3 Scratchpad files and /compact: keep the findings, drop the noise
If the conversation can't be trusted to hold the morning's findings, they have to live somewhere else. Back in the contract room, a sensible person jots "contract 17, clause 4" on a pad the moment they find it. The agent's version is a scratchpad file: a plain file where the agent records each key finding as it discovers it, and which it reads back before answering later questions. It is a pattern, not a product feature. Any ordinary file works; here it is called findings.md.
Because the file sits outside the context window, it survives every context boundary, the points where a conversation is summarised, emptied or lost. Compaction, a new session, a subagent's fresh start and a crash are all boundaries. When the engineer asks about stock at four o'clock, the agent reads findings.md first and answers from an exact copy of the morning's findings, not a buried one.
Two habits make it work: the agent writes to the file the moment it finds something, and reads it before answering any question about the system. Put both rules in the agent's system prompt, its standing instructions, or in CLAUDE.md, the project instructions file the agent loads at the start of every session. Don't leave them in an early chat message: compaction can drop early instructions, while CLAUDE.md is re-injected on every request.
In the file below, look at the path on each class line and the open question at the end. The paths keep later answers specific, and the open questions tell a later session where to pick up.
# Order-fulfilment system: key findings
## Classes (name | file | role)
- FulfilmentOrchestrator | orders/fulfil.py | entry point for every order
- WarehouseAllocator | stock/alloc_v2.py | reserves stock; alloc_v1.py is dead code
- RefundLedgerService | billing/rl_svc.py | post_reversal() writes refunds to the ledger
- LegacyOrderDAO | db/legacy_dao.py | every order read and write goes through it
## Facts established
- A refund never calls WarehouseAllocator; restocking is a nightly job, jobs/restock.py
## Open questions
- What else calls post_reversal() besides the admin screens?
With the findings safe on disk, the search output in the conversation is dead weight. That is the moment for /compact, a command in Claude Code, Anthropic's coding assistant for the terminal. It summarises the conversation so far and carries on from the summary. You can say what to keep, as in /compact Focus on class names, file paths and the refund flow, or put standing compaction instructions in a section of CLAUDE.md. Your SDK-based tool can trigger the same thing by sending /compact as an ordinary prompt.
Order matters. Automatic compaction runs near the limit whether you are ready or not, and keeps what it guesses is important. Running /compact yourself, right after the findings are written down, lets you choose the moment and the focus. The scratchpad guarantees the specifics; /compact clears the noise around them.
5.4.4 Subagents: one question each, while the main agent keeps the map
A scratchpad keeps findings safe, but it doesn't stop the window refilling. Suppose the main agent runs every search itself. One question such as "find all test files" can mean hundreds of file names from Glob and dozens of file peeks, all landing in the context meant to hold the big picture. A subagent moves that reading elsewhere. It is a separate agent with its own fresh, isolated context window. The main agent hands it a task, it reads whatever it needs, and only its final answer comes back, the same way any tool's output does.
Think of a newspaper editor on deadline. The editor doesn't interview forty people; they send one reporter per story, keep the shape of the whole edition in their head, and get back finished paragraphs rather than raw interview notes. The main agent likewise keeps high-level coordination: which modules exist, how they connect, which questions are open, and what to ask next. Its context grows by the answers, not by the hundreds of hits behind them. When several agents work together like this, the main agent is usually called the coordinator.
The briefs must be specific. Each is one question with a clear finish line: "find all test files", "trace the refund flow dependencies", "list every class that calls LegacyOrderDAO directly". Each also names the return it wants, such as a list of paths, or a call chain (which function calls which) with file names. A brief like "explore the codebase" only moves the degradation into the subagent, which returns an essay that refills the main context. Even good briefs add up, because many subagents returning detailed results can consume significant context. So ask for short, structured answers and record them in the scratchpad.
One map keeper, one question per subagent
LegacyOrderDAO?"In Claude Code, the built-in Explore subagent is a ready-made, read-only version of this pattern. Keep one property in mind for the next section: a subagent normally starts without the main agent's conversation, so anything it needs from earlier findings must be in its prompt.
5.4.5 Between phases: summarise, then brief the next wave
Large explorations come in phases. Phase one maps the order-fulfilment system: modules, entry points, the classes everything depends on. Phase two digs into specific risks, such as "which code paths can post the same refund twice?" and "which tests cover allocation?". A phase-two subagent sent off with only its question rediscovers LegacyOrderDAO from scratch, spending its own context to do it. With nothing to anchor it, it may describe the classes differently or reach for typical patterns.
The tempting overcorrection is to paste phase one's whole transcript into each new prompt. That imports the very output you delegated to keep out, copied into every subagent. The right move sits between nothing and everything. At the end of each phase, the main agent summarises the key findings: classes with paths, how they connect, decisions made, open questions. It then injects that summary into the initial context of every next-phase subagent, alongside that subagent's one question. Injecting just means the coordinator puts the summary text at the top of each prompt it writes; no feature does it for you.
A hospital shift handover works the same way. The night nurse gets neither a recording of the day shift nor nothing, but a short handover: this patient, this medication, watch for this. The phase summary is that handover, written from the scratchpad and short enough for every subagent to carry.
Summarise one phase before starting the next
Here is the start of the refund subagent's prompt: five lines of injected summary, then its own question and the shape of the answer it should return.
What phase 1 established:
- Orders enter through FulfilmentOrchestrator (orders/fulfil.py).
- Every order read and write goes through LegacyOrderDAO (db/legacy_dao.py).
- Refunds post through RefundLedgerService.post_reversal() (billing/rl_svc.py).
- Restocking is a nightly job (jobs/restock.py), not part of the refund call.
Your question: which code paths can call post_reversal() twice for one order?
Return: each path as a call chain with file names, under 200 words.
5.4.6 Crash recovery: state exports and a manifest
By now the exploration is a multi-agent run: a coordinator and eight subagents working through phase two. At three o'clock the process dies: perhaps a laptop goes to sleep, a server restarts or the machine runs out of memory. Whatever lived only in the agents' conversations is gone. Rerunning from the start repeats hours of work that may not even reach the same conclusions.
The design that survives this is structured state persistence. Each agent exports its state to a known location, a path fixed before the run starts, such as state/refund-tracer.json. (JSON is a plain-text format of labelled fields that programs read easily.) The agent writes at checkpoints rather than only at the end: what it has established, which files it has covered, what it would do next. The coordinator keeps a manifest, one file listing every task with its status and the path to its state file, plus the latest phase summary.
Like the scratchpad, this is a pattern you design, not a product feature: you choose the files, the fields and the code that reads them. Your code updates the manifest when a subagent starts or finishes. The Agent SDK can call a function of yours, a hook, at exactly those moments: the SubagentStart and SubagentStop hooks. In the manifest below, status tells the coordinator what to do with each task, and state says where that agent's saved progress lives.
{
"run": "order-fulfilment-exploration",
"phase": 2,
"phase_summary": "state/phase1-summary.md",
"tasks": [
{"question": "find all test files",
"status": "done", "state": "state/test-finder.json"},
{"question": "trace the refund flow dependencies",
"status": "in_progress", "state": "state/refund-tracer.json"},
{"question": "which tests cover allocation?",
"status": "pending", "state": null}
]
}
On resume, the coordinator loads the manifest before anything else. Think of a save point in a long video game: when the power cuts out, you reload at the last save with your inventory intact, not at the title screen. Then the coordinator acts on each task's status.
status |
What it means | What the coordinator does on resume |
|---|---|---|
done |
The agent finished; its results sit in its state file | Skips the task and reads the results from the file |
in_progress |
The agent was mid-task when the crash hit | Respawns it with its saved state injected into its prompt |
pending |
The task never started | Spawns it fresh, with the phase summary and its question |
The status names are this example's choice; the rule behind them is what to remember: never redo finished work, and resume unfinished work from its saved state. A respawned refund tracer's prompt might begin: "You are resuming 'trace the refund flow dependencies'. Already established: ... Files covered: ... Continue from: callers of post_reversal() in jobs/."
The location is fixed in advance because after a crash there is nobody to ask where things were saved. The state is structured because "skip or resume" must be decided reliably, and a status field is a fact your code can check, where a paragraph of prose is not. Resuming the crashed conversation itself, where that is possible, would bring back the whole noisy history; the manifest brings back only what each agent had established.
Recovering from a crash with a manifest
5.4.7 The exam traps
Each trap below either skips the right technique or uses it badly.
- ✗ Treating vague afternoon answers as a prompting problem ("always name specific classes"). ✓ Recognise context degradation and move the findings into a scratchpad file the agent writes on discovery and reads before answering. An instruction cannot restore details the model can no longer reliably see.
- ✗ Letting the main agent read everything itself so that "nothing is lost in handoffs". ✓ Delegate each specific question to a subagent. The verbose output stays in the subagent's context, and the main agent keeps its room for the map.
- ✗ Giving one subagent a broad brief such as "explore the whole codebase". ✓ One narrow question per subagent, with a short, defined return. A broad brief moves the degradation into the subagent and sends an essay back.
- ✗ Starting next-phase subagents with only their question, or with the full phase-one transcript. ✓ Inject a summary of the phase's key findings into their initial prompts. With nothing, they rediscover or guess; with everything, they re-import the noise.
- ✗ Restarting a crashed multi-agent run from the beginning. ✓ Have each agent export state to a known location and have the coordinator load a manifest on resume, injecting saved state into the unfinished agents' prompts.
- ✗ Running
/compact, or clearing the session, before the key findings are written down. ✓ Write the findings to the scratchpad first, then/compactwith a focus. Compaction may drop specifics, and a cleared session keeps only what is on disk.
5.4.8 Put it together: explore a system and survive a crash
You now have the whole toolkit. Each technique answers a different problem, and exam questions describe the problem, so learn them as pairs.
| The problem you see | The technique | Why it works |
|---|---|---|
| Answers drift to "typical patterns" and contradict earlier ones | Scratchpad file of key findings, read before answering | Findings live on disk, outside the crowded window |
| Each new question floods the main agent's context with search hits and file reads | Subagents, one specific question each | Verbose output stays in the subagent; a short answer returns |
| Next-phase subagents rediscover old findings or guess | Summary of the last phase injected into their initial prompts | They start from what is known, without the transcript |
| A long multi-agent run crashes halfway | State exports plus a manifest the coordinator loads on resume | Finished work is skipped; unfinished agents resume from saved state |
| The session is already full of old discovery output | /compact with a focus, after the findings are saved |
Summarises the conversation and frees the space |
Memorise the first two columns as pairs; the third is the reason you need when two options look alike.
The rest of the domain turns from keeping findings to trusting them. Human review and confidence calibration (5.5) decide which agent outputs a person should check before anyone acts on them. Provenance (5.6) asks every claim to keep its source attached, which your scratchpad already does when each finding carries its file path.
Key takeaways
- ✓ Context degradation shows as an agent citing "typical patterns" instead of the specific classes it found, and contradicting its own earlier answers.
- ✓ A scratchpad file, an ordinary file the agent is told to keep, holds key findings with their file paths outside the conversation; the agent writes on discovery and reads it before answering later questions.
- ✓
/compactsummarises the conversation to clear verbose discovery output; save the findings first and give it a focus on what to keep. - ✓ Subagents each investigate one specific question in their own context and return a short answer, while the main agent keeps the high-level map.
- ✓ Between phases, summarise the key findings and inject the summary into the next subagents' initial prompts: never nothing, never the whole transcript.
- ✓ For crash recovery, agents export structured state to a known location, and the coordinator loads a manifest you designed on resume and injects each unfinished agent's state into its prompt.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
54 CCAR-F questions on Domain 5, free
Every question in the bank is tagged to a domain, so you can drill 54 questions on Context Management & Reliability alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 5 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.