Home › Study guides › CCAR-P › Domain 3 › Lesson 3.8
CCAR-P · Domain 3 · 19% of the exam · Lesson 3.8 · 21 min read
Progressive discovery or monolithic context: what an agent loads up front
When to load every tool, runbook and document up front, when to let the agent discover detail on demand, and how to design the hybrid in between.
Written against objective 3.8 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.8.1 Why the IT copilot got worse as it learned more
Ambergate Software builds payroll products. Eighteen months ago its IT team launched an internal operations copilot for the service desk: a dozen tools for the ticketing system and the identity provider, twenty runbooks, everything in the prompt. It worked, so every team added to it. Today it holds 60 tools across ticketing, identity, cloud and network systems, and 120 runbooks.
Every request still carries all of it: the 60 tool definitions (about 31,000 tokens), the operations policy, and a summary of each runbook, since the full runbooks stopped fitting comfortably long ago. That is about 60,000 tokens before an engineer types a word. Prompt caching, which lets the API reuse an unchanged prompt prefix at a fraction of the normal input price, keeps the bill tolerable. What it cannot fix shows up in the traces. The copilot calls idp_reset_password when the request needed idp_reset_mfa, and it answers from a two-line runbook summary when the step that mattered was on page three of the real runbook.
Halvard, who runs IT operations, asks Zainab, the solution architect, for a stronger model. She sees a context strategy problem instead. The copilot uses a monolithic context strategy: it loads every instruction, tool definition and piece of reference material up front, on every request. The alternative is progressive discovery: load a short description of what exists, and let the agent pull in the detail when a task needs it. Neither is right in general; most production systems end up with a hybrid.
What one request carries
Monolithic
Progressive discovery
3.8.2 Monolithic context: simple, predictable and heavy
Why did Ambergate start monolithic? Because for a dozen tools and twenty runbooks it was the right call, as it still is for many systems.
A monolithic context has nothing to discover, so nothing the agent needs is ever left unloaded. Every request sees the same instructions in the same order, which makes behaviour easy to reproduce: replay a request and the model sees exactly what it saw in production. The unchanging prefix is ideal for prompt caching. There are no extra steps before the agent can act, and there is one prompt to build, review and test.
The costs grow with the catalogue, and they compound. Tokens come first: every request pays for every tool and runbook, used or not. Caching makes those tokens cheaper to resend; it does not make them lighter to reason over. Attention comes next. Every irrelevant definition competes with the relevant one, and Anthropic's tool search documentation notes that Claude's ability to pick the right tool degrades once more than 30 to 50 tools are available. Look-alike names such as idp_reset_password and idp_reset_mfa make it worse.
Then the limits arrive. When the catalogue no longer fits comfortably, teams compress it, and summaries replace documents. Ambergate's runbook summaries are that compromise, and they lose exactly the detail an engineer needs at 2 a.m. Finally, every addition changes the prompt that every request receives, so one new runbook is a change to the whole system.
Why Ambergate's monolith stopped scaling
3.8.3 Progressive discovery: an index first, detail on demand
How does an experienced on-call engineer handle 120 runbooks? Not by memorising them. They know which runbooks exist, roughly what each is for, and how to open the right one in seconds. Progressive discovery gives an agent the same arrangement: short descriptions stay loaded, and the agent fetches only the detail its task needs.
Anthropic's context engineering guidance calls the fetching half just-in-time context: the agent holds lightweight identifiers, such as file paths, stored queries or links, and loads the data at runtime with tools. Claude Code is Anthropic's own example of mixing up-front and just-in-time context: it drops CLAUDE.md files into context up front and uses glob and grep to find files just in time. The pattern appears in four mechanisms an architect can combine.
| Mechanism | What stays loaded | What loads on demand, and who decides |
|---|---|---|
| Tool search with deferred loading | The search tool and any tools not marked defer_loading |
Full definitions of the tools a search matches; Claude searches, the API expands the matches |
| Agent Skills | Each Skill's name and description | The SKILL.md instructions, then bundled files; Claude decides when a task matches the description |
| MCP resources | The resource list (a URI, name and description each), held by your application | The contents read with resources/read; the host application decides, because resources are application-driven |
| Just-in-time retrieval | Identifiers and a way to use them: an index, IDs, a search tool | The document or record looked up; Claude decides, through tools you provide |
Read the last column carefully. In three mechanisms the model decides what loads, so you evaluate its discovery judgement. MCP resources differ: MCP tools are model-controlled, but resources are application-driven. Your host application decides how they reach Claude: it picks them itself, lets the user attach them, or gives Claude tools to list and read them, as Claude Code does. On the Messages API, the MCP connector carries MCP tools only, so a design built on resources needs an MCP client in your own application.
Zainab maps Ambergate's catalogue onto the table. The 60 tools suit tool search. The runbooks suit just-in-time retrieval: a one-line index stays loaded, and a read_runbook tool fetches the full text by its ID when the agent needs it. A resource list chosen by the application would not do, because only the agent knows, halfway through a task, which runbook it needs next.
One request under progressive discovery
idp_reset_mfaread_runbook("IDN-014"), full text3.8.4 What discovery costs: steps, misses, caching and traces
It is tempting to treat progressive discovery as free savings. Resist it: it trades tokens and attention for four new costs, in steps, misses, cache design and traceability.
Steps and latency. A tool search adds a step before the tool call; it runs on Anthropic's servers within the same response, but the model still writes the query and reads the results. Every read_runbook is a full round trip: the model asks, your code fetches, the model reads. Anthropic's context engineering guidance puts it plainly: exploring at runtime is slower than having the data ready. Discovery pays only when the savings in tokens and the gains in accuracy outweigh those extra steps.
The miss. A monolith never has to find a tool. A discovery design does, and it can fail. The descriptions do the routing. A vague name, or an engineer's wording ("VPN won't connect") that shares no words with the catalogue's ("remote access gateway"), can leave the agent concluding that no suitable tool exists. The docs' mitigations are design work: clear names and descriptions in the words users use, prefixes that group tools by system, a system-prompt line naming the tool families, and the most-used tools kept loaded.
Caching. A monolith gives one stable prefix. Discovery keeps caching only if what it loads lands after that prefix. Tool search is built for this: deferred tools are left out of the prefix, and discovered ones are appended inside the conversation. A home-grown router that rewrites the tools array for each request does the opposite, because changing tool definitions invalidates the entire cache: tools, system prompt and messages.
Predictability and observability. Unlike a monolith, a discovery design assembles its context per request, so two runs of one question can load different tools. That is why traces must record the discovery steps: each search query, what it returned, each runbook fetched and each search that found nothing. Your evals need a new measure too, discovery recall: for each test case, did the agent load the tool or runbook the case needs?
| Concern | Monolithic | Progressive discovery |
|---|---|---|
| Tokens per request | The whole catalogue, every time | A map plus what the task loads |
| Selection as the catalogue grows | Degrades as tools and documents compete | Stays focused on a few candidates |
| Steps before acting | None | A search or fetch per missing piece |
| Risk of never loading what is needed | None: everything is loaded | Real; depends on descriptions and search |
| Prompt caching | One stable prefix | Preserved only if loads land after the prefix |
| Reproducing a request | Replay it | Needs the logged discovery steps |
3.8.5 Four questions that decide the strategy
"Should we use progressive discovery?" has no answer until you describe the catalogue. Four questions settle most cases, answered from traces and content owners, not preference.
- SIZE AND VARIETY. How big is the catalogue, and how different are its parts? The larger and more varied it is, the smaller the slice any one request needs. The tool search docs suggest discovery from about ten tools or 10,000 tokens of definitions, and plain tool calling below ten.
- USAGE. How often is each part used? Most catalogues are skewed: a small head appears in most requests and a long tail appears rarely. Load the head; discover the tail.
- STABILITY. How often does it change, and who owns it? Stable content is cheap to cache and safe to load. Content that changes weekly, or lives in another team's system, is better fetched at request time, so the agent reads the current version and an edit never rewrites the prefix.
- LATENCY BUDGET. How many extra steps can a request afford? A nightly job can take several discovery steps; an interactive assistant with a tight p95, the 95th-percentile response time, may afford one.
One rule overrides all four. Content that must govern every request, such as a change freeze, an approval rule or a safety constraint, stays loaded whatever its usage. A rule that has to be discovered is a rule that can be missed. Loading it makes Claude aware of it on every request; where a breach would do real harm, the system behind the tool enforces it as well.
| Option | When it wins | What it costs |
|---|---|---|
| Monolithic, cached | Small, stable catalogue used on most requests; tight latency; identical context wanted for every request | Tokens and attention on every request; selection degrades as it grows; every change touches every request |
| Progressive discovery | Large, varied or fast-changing catalogue, each request needing a small part; the latency budget allows extra steps | Extra steps; the risk of a miss; descriptions to maintain; discovery to log and evaluate |
| Hybrid | Skewed usage: a small head used on most requests, rules that must always apply, and a long tail | Two layers to design, and a line between them to revisit as usage shifts |
Claude Code ships these choices as settings. By default it loads only MCP tool names at session start and finds full definitions through tool search; a server marked alwaysLoad: true joins the always-loaded core. With ENABLE_TOOL_SEARCH=auto, it loads MCP tools up front only while their definitions stay under 10% of the context window.
3.8.6 The hybrid: an always-loaded core, a discoverable long tail
Where should Ambergate's line fall? Zainab reads 30 days of traces. Four tools appear in most requests: fetching, updating and searching tickets, and looking up a user. The other 56 form a long tail, some used weekly and some only at quarter end. Runbooks live in the wiki, where four teams edit them every few weeks. The operations policy, with its change freeze and approval rules, must govern every request.
The core stays loaded and cached: the policy, the four head tools, a read_runbook tool, a line naming the tool families, and a runbook index of one line each (ID, title, when to use it). The index changes only when a runbook is added or retired, so routine edits never touch the cached prefix.
The tail is discoverable: Zainab gives the other 56 tools system prefixes (itsm_, idp_, cloud_, net_), rewrites their descriptions in the words engineers use, and defers them behind tool search. Full runbooks come from the wiki through read_runbook, so an edit reaches the next request. Each role (service desk, cloud, network) gets its own fixed tools array and index, listing only what it may use, and the systems behind each tool still check permissions when it runs.
Zainab also considered packaging the runbooks as Skills. She rejected it because each wiki edit would then need a repackaging and upload step, and the wiki would stop being the single source. In the request below, look at defer_loading on the tail and at the one cache breakpoint on the system prompt: because the cache prefix runs from tools to system, it covers the core tools too.
CORE = [itsm_get_ticket, itsm_update_ticket, itsm_search_tickets,
idp_get_user, read_runbook] # the head: always loaded
TAIL = [dict(t, defer_loading=True) for t in tail_tools] # 56 tools, found by search
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
system=[{
"type": "text",
"text": OPS_POLICY + TOOL_FAMILIES + RUNBOOK_INDEX, # rules and the map
"cache_control": {"type": "ephemeral"}, # caches tools + system: the stable prefix
}],
tools=[{"type": "tool_search_tool_bm25_20251119", "name": "tool_search_tool_bm25"}]
+ CORE + TAIL, # fixed per role, never per request
messages=messages,
)
On a replay of 400 past requests, tokens before the question fell from about 60,000 to about 10,000, and wrong-tool calls fell from one in eight to one in thirty. Median response time barely moved, because most requests never left the core. Requests that needed a tail tool paid one extra search step, and that showed at p95. Her decision record names the price too: discovery recall joins the eval suite, and she reviews the line every quarter.
Ambergate's hybrid context
itsm_ tools, deferredidp_ tools, deferredcloud_, net_ tools, deferredread_runbook3.8.7 The exam traps
Most traps here either keep the monolith and patch around it, or adopt discovery without paying its costs.
- ✗ Keeping everything loaded and moving to a bigger window or a bigger model because tool choice is erratic. ✓ Move the long tail behind discovery. A bigger window postpones the limit without restoring focus. A larger model still sees every competing tool.
- ✗ Compressing the catalogue into summaries so the monolith fits. ✓ Keep the full detail at its source and load it when needed. Summaries lose the step that matters and go stale.
- ✗ Making everything discoverable, including rules every request must follow. ✓ Always load the must-apply rules and the head the traces show. Discovery can miss.
- ✗ Rebuilding the
toolslist per request with your own router. ✓ Keep one stable prefix and load detail after it, as tool search does. A changed tools array invalidates the whole cache. A router's wrong guess also hides a tool completely. - ✗ Adopting discovery for a small catalogue that every request uses. ✓ Send it whole and cache it. A search step buys nothing when every tool is needed anyway.
- ✗ Logging only the final answer. ✓ Log each discovery step and measure discovery recall. Without the trace, a search that found nothing looks like a model mistake.
3.8.8 Put it together: design and test a hybrid context
You now have every piece: both strategies and what they cost, the four discovery mechanisms, the four deciding questions and the hybrid. The quickest way to make it stick is to measure both strategies on one small catalogue, then break discovery on purpose.
Later domains build on these runs. Optimising token usage, latency and cost (4.5) weighs the tokens you saved against the steps you added, and monitoring with logging and observability tools (4.6) turns discovery traces and recall into dashboards and alerts. Configuring Claude tools for teams (7.1) sets the same line in Claude Code for a whole organisation.
Key takeaways
- ✓ Monolithic context loads everything on every request: predictable, simple, reproducible and cache-friendly, but heavy, diluting and hard to grow.
- ✓ Progressive discovery keeps short descriptions loaded and fetches detail on demand, through tool search, Skills, MCP resources or just-in-time retrieval.
- ✓ The model decides what loads with tool search, Skills and retrieval tools; with MCP resources, the host application decides how they reach Claude.
- ✓ Discovery costs extra steps, the risk of a miss and less predictable behaviour, so log every discovery step and measure discovery recall.
- ✓ Keep the cached prefix stable and load detail after it; rewriting the tools array per request invalidates the whole cache.
- ✓ Decide from size and variety, usage, stability and latency budget, and always load rules that must govern every request.
- ✓ Most catalogues are skewed, so the usual answer is a hybrid: a cached core of rules, head tools and a map, plus a discoverable long tail.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAR-P questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Integration alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.