Home › Study guides › CCAR-P › Domain 3 › Lesson 3.6
CCAR-P · Domain 3 · 19% of the exam · Lesson 3.6 · 22 min read
Retrieval strategies matched to the data and the question
Dense, keyword and hybrid search, filters, SQL tools and agentic search: which fits which data and question, why vectors cannot count, when to skip it.
Written against objective 3.6 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
3.6.1 Why one search index cannot answer every question
Tomasz, a field-service engineer at Veldmark Compressors, is standing at a tripped compressor in a bottling plant, and the panel shows E-417. He has three questions for the assistant on his tablet. What does E-417 mean? Which shaft seal fits this machine, known to the customer as unit 7B? And how many tickets report overheating since the 2024 firmware release? The answers live in four places: 2,000 pages of service manuals, a parts database (a 14,000-row catalogue plus a register of every installed unit), 90,000 past service tickets and a table of 900 fault codes.
The pilot cut all four sources into short passages, called chunks, and put them in one vector index, a search index that matches text by meaning. For every question it sent Claude the ten most similar chunks. Claude answered all three with confidence and got all three wrong. It explained E-471, a different fault. It recommended a seal kit that a 2023 service bulletin had replaced. And it reported that 10 tickets mention overheating, because it had retrieved 10 tickets; 212 in the ticket database match.
Nothing was wrong with the model. Each question went to the same kind of search, and that search suits only one kind of question. Solveig, the architect Veldmark brings in, treats retrieval as a choice, not a component. A retrieval strategy is how the system finds what the model needs: which index or system it asks, how it matches, and how much it brings back. It is chosen per kind of question, from two inputs: the shape of the data and the pattern of the query.
One index for everything versus a strategy per question
One vector index
A strategy per question
3.6.2 Name the data shape and the query pattern first
"Which retrieval method is best?" has no more of an answer than "which database is best?". An architect asks a narrower question: best for which data, asked in which way? Answer both before anyone mentions an index.
The first input is the data shape, how the information is stored and how meaning sits in it:
- Long prose (manuals, policies): meaning is spread across paragraphs, and the same idea can be phrased many ways.
- Semi-structured records (tickets, emails): a few reliable fields, such as unit, date and firmware, plus free text. The fields can be filtered and counted; the text can only be searched.
- Structured tables and databases (the parts catalogue): rows with exact values, where the answer is a query, not a passage.
- Code: exact names such as functions and configuration keys, plus structure such as files and calls that an answer often has to follow.
- Codes and identifiers (fault codes, part numbers, serials): short strings where one character changes the meaning, so a near match is a wrong match.
The second input is the query pattern, what a correct answer requires. Six patterns cover almost everything Tomasz and his colleagues ask.
| Query pattern | What a correct answer needs | At Veldmark |
|---|---|---|
| Exact lookup | The one record matching a key, character for character | "What does E-417 mean?" |
| Semantic question | Passages that mean the same as the question, however worded | "Why would a dryer short-cycle in cold weather?" |
| Multi-hop | Facts from several places, each found using the previous one | "Which seal fits unit 7B?" |
| Aggregation | A computation over every matching record, not a sample | "How many tickets mention overheating after the 2024 firmware?" |
| Recency-sensitive | The newest valid version, not the most similar one | "What is the latest bulletin on oil carryover?" |
| Filtered by attribute | Only items that meet a hard constraint | "Past fixes for this fault, on this model only" |
Learn the six patterns so you can spot them in a question stem. Real questions often combine two or three: the ticket count is an aggregation, filtered by firmware, over a symptom that engineers describe in many different words.
3.6.3 Meaning or exact tokens: dense, keyword and hybrid search
Why did the pilot explain the wrong fault code? Because of how dense vector search works. An embedding model turns each chunk, and at query time the question, into a vector, a list of numbers that encodes meaning, and the search returns the chunks nearest the question. That makes it superb for semantic questions: "running hot", "thermal trip" and "overheating" land close together.
The same property is the weakness. An embedding captures meaning, not characters, so E-417, E-471 and E-414 all mean roughly "a compressor fault code". Anthropic's contextual retrieval write-up gives the same example: an embedding model may find content about error codes in general and miss the exact code.
Keyword search fixes this. Its usual form is BM25 (Best Matching 25), a ranking function that scores text by the exact terms it shares with the query and weights rare terms heavily. E-417 is rare, so its own passage ranks first. The weakness mirrors the strength: a ticket that says "running hot" never matches a query for "overheating".
Hybrid search sends the query to both and merges the two ranked lists into one, removing duplicates (a step called rank fusion). Re-ranking then scores each candidate against the question with a separate reranking model and keeps the best few. Think of two librarians, one who knows the subject and one who knows the catalogue numbers, with an editor choosing the best twenty from both piles. Anthropic's retrieval tests found that both steps pay: embeddings plus BM25 beat embeddings alone, and re-ranking beat none. The re-ranker runs on every query, so how many candidates you give it trades accuracy against latency and cost.
Hybrid search with re-ranking
Solveig puts the manuals and the free text of the tickets behind hybrid search with re-ranking. The fault-code table gets no search at all: when a question contains a code, a lookup by key returns that one row, character for character. Chunking and indexing are a separate decision; here the indexes exist, and the question is how to query them.
3.6.4 Counts, attributes and dates: filters and structured queries
Here is the failure that surprises teams most. Similarity search returns the top k most similar chunks, where k is a number you set. By design it is a sample of the best matches, never all matches. If k is 10, the largest count Claude can possibly see is 10. The pipeline asked a sampling tool for a census.
Aggregation (how many, which top five, what trend) needs a structured query: a query that runs in the database over every matching record, reached through a tool that Claude calls. Claude turns the question into parameters and explains the number that comes back. Your application runs a fixed, parameterised, read-only query. If you let the model write free SQL instead, you gain flexibility and must add a read-only database role, row limits and validation of every generated query.
Counting from search results versus counting with a query
Count from retrieved chunks
Count with a query
count_ticketssymptom, firmware, datesOne snag: "mention overheating" is free text, not a field. Solveig has Claude tag each ticket with a symptom category, in batch at ingestion, so the count runs on a field. A full-text count with a synonym list is cheaper and misses every phrasing nobody listed. Look at two things in the tool she specifies: the description says when to use it and when not, as Anthropic's tool guidance advises, and symptom is an enum over that tagged field.
count_tickets = {
"name": "count_tickets",
"description": (
"Counts ALL service tickets that match every filter and returns the exact count "
"plus up to 5 example ticket IDs. Use it for 'how many', 'which most' and trend "
"questions about tickets. Do not use it to read what tickets say; use "
"search_tickets for that."
),
"input_schema": {
"type": "object",
"properties": {
"symptom": {"type": "string", # tagged at ingestion, so the count runs on a field
"enum": ["overheating", "oil_carryover", "vibration", "pressure_loss"]},
"model": {"type": "string", "description": "Compressor model, e.g. VX-90"},
"firmware_from": {"type": "string", "description": "Lowest firmware version, e.g. 5.0"},
"opened_after": {"type": "string", "description": "ISO date, e.g. 2024-03-01"},
},
"required": ["symptom"],
},
}
Hard constraints get the same treatment. Metadata filtering applies conditions such as model, site, document status or effective date in the retrieval layer, before ranking. The alternative, telling Claude to ignore passages about other models, fails twice: the wrong passages already took slots in the top k and tokens in the prompt. Recency works the same way. Similarity has no idea what is newest, and a 2019 bulletin can read closer to the question than the 2023 one that replaced it. Filter on status and date, or sort relevant results by date.
3.6.5 Multi-hop questions: agentic search and query rewriting
"Which seal fits unit 7B?" No chunk anywhere contains both "7B" and the answer. The unit register says 7B at this plant is a VX-90 built in 2019. The parts catalogue says VX-90 units built before 2020 take seal kit SK-2214. A service bulletin in the manuals says SK-2231 replaced SK-2214, and only the newest bulletin counts. Each search needs the answer to the one before it, so a pipeline that retrieves once, up front, cannot answer the question at all.
Agentic search gives Claude search tools and lets it run them in a loop: search, read the results, decide the next search, stop when it has the answer. Claude requests each search; your application runs it and enforces a cap on tool calls. Claude Code works this way on a codebase. Rather than querying a pre-built search index, it runs glob and grep searches over file names and contents, reads what it found and searches again. Anthropic describes this just-in-time approach as sidestepping stale indexes, which makes it a strong fit for code and for sources that change faster than you could re-index them.
Anthropic's context-engineering guidance names the costs: exploring at runtime is slower than retrieving pre-computed data, and an agent without good tools and guidance wastes context on dead ends. Its middle ground: retrieve some data up front for speed and let the agent explore further only when it needs to. And the loop earns its turns only when each hop depends on the last. When the sources are known in advance and independent, such as the manual and the ticket history for one fault code, your code queries them in parallel and calls Claude once.
Three hops to one answer
lookup_unit: a VX-90 built in 2019query_parts: VX-90 before 2020 takes SK-2214search_manuals: a bulletin replaces it with SK-2231Query rewriting turns the user's words into what an index can match. "The 2024 firmware" becomes version 5.0, "tripping hot" becomes the symptom overheating, a two-part question becomes two sub-queries, and a follow-up such as "and on the older model?" becomes a complete question. Inside agentic search Claude writes its own queries, so rewriting comes built in. In a fixed pipeline, a call to a small, fast model before retrieval does the job, for one extra call's latency. It earns its place wherever users and documents use different words.
3.6.6 The matching table, and when not to retrieve at all
The design is now a routing decision. Patterns that code can detect, such as E- followed by three digits, are routed in code: deterministic and cheap. For open questions Claude chooses among a few clearly distinct tools, because overlapping retrievers with vague descriptions make that choice unreliable.
| Strategy | When it wins (data shape; query pattern) | What it costs or where it fails |
|---|---|---|
| No retrieval: corpus in a cached prompt | A small, stable corpus used on most requests; or a question that needs no company data | Tokens on every request (much cheaper once cached); cost and recall suffer as the corpus grows |
| Direct lookup by key | Identifiers and structured records; an exact lookup the router can detect | Needs a clean key and an index or API; no help with paraphrase |
| Keyword search (BM25) | Codes and rare terms inside text; exact lookup | Misses synonyms and paraphrase |
| Dense vector search | Long prose; semantic questions | Misses exact identifiers; cannot count, find the newest or guarantee completeness |
| Hybrid with re-ranking | Prose and the free text of records; mixed traffic, the usual default | Two indexes to maintain; the re-ranker adds latency and cost |
| Metadata filter | Records and documents with clean fields; attribute filters and recency | Needs clean attributes on every item; a wrong filter hides good results |
| Structured query tool (SQL or API) | Tables and record fields; aggregation: counts, top N, trends, joins | Needs fields to query, extracted at ingestion if necessary; an interface to build and secure |
| Agentic search | Code, linked sources, fast-changing sources; multi-hop and exploratory questions | More turns: slower, more tokens, less predictable latency |
| Query rewriting (added to any of the above) | User wording that differs from the documents; follow-ups; compound questions | One extra model call; a careless rewrite can drop the user's intent |
Memorise two facts above all: vector search cannot count, and it cannot reliably match exact identifiers. Recognise the rest.
How much to retrieve is its own trade-off. Too few passages and the answer may be missing; too many and you pay in tokens and latency, and a model's recall of details drops as its context fills. In Anthropic's retrieval tests, passing 20 chunks beat 10 and 5, with the advice to test on your own data. The usual winner: retrieve wide, re-rank, send few. Tool results follow the same rule, which is why count_tickets returns a number and five example IDs, not 212 tickets.
Sometimes the best retrieval is none. When a corpus is small, stable and needed on most requests, put all of it in the prompt, ahead of the question, and cache it. Anthropic's contextual retrieval write-up puts that line at roughly 200,000 tokens, about 500 pages, and cached reads are billed at a small fraction of the normal input price. Questions that need no company data, such as a unit conversion or a follow-up on what the conversation already holds, need no search either. A pipeline that retrieves on every turn pays for retrieval on "thanks, that worked".
Here is the routing record Solveig signs; look at the last two lines, which say what is NOT retrieved and what evidence backs each route.
Retrieval routing, field-service assistant, v1. Owner: Solveig (architect).
Fault codes ("what does E-417 mean"): code detects the pattern and calls the fault-code lookup; keyword search adds the manual's troubleshooting section for that exact code. No vector search.
Manual questions: hybrid search with re-ranking, filtered to the unit's model and current manual revisions; about 150 candidates re-ranked to 20.
Parts and unit questions ("which seal fits unit 7B"): agentic search over lookup_unit, query_parts and search_manuals, capped at 6 tool calls.
Ticket counts and rankings: count_tickets on the ticket database; Claude explains the number and cites up to 5 example tickets. Similar past cases: hybrid search on ticket notes, filtered by model, newest first.
Not retrieved: the lockout and safety procedure (in the cached system prompt); conversions and follow-ups answered from the conversation.
Evidence: 120 real engineer questions with known answers, scored per route before launch and after every index change.
3.6.7 The exam traps
Almost every trap here treats retrieval as one component to tune rather than a choice to match.
- ✗ Sending every source and every question to one vector index. ✓ Route by data shape and query pattern; one similarity search fails silently on codes, counts and chains of facts.
- ✗ Letting Claude count or rank from retrieved chunks. ✓ Run aggregations as structured queries over every record and have Claude explain the result. A top-k sample caps any count at k.
- ✗ Fixing identifier misses with more chunks, a bigger model or a prompt to "check the code carefully". ✓ Add keyword or hybrid search, or a direct lookup. If the right passage was never retrieved, nothing downstream can recover it.
- ✗ Replacing dense search with keyword search once identifiers fail. ✓ Keep both in hybrid search. Keyword alone breaks the paraphrased questions that already work.
- ✗ Asking Claude to ignore passages about the wrong model or an old version. ✓ Apply metadata filters in the retrieval layer before ranking, so ineligible passages never compete for the top k.
- ✗ Using agentic or multi-agent search for every question, or retrieving on every turn. ✓ Match the machinery to the question: a lookup or fixed pipeline for predictable questions, agentic search only where hops depend on each other, and no retrieval when no company data is needed.
3.6.8 Put it together: route each question to the search that fits
You now have every piece of the decision. The quickest way to make it stick is to watch similarity search fail at a count and at an exact code, then fix each.
Later objectives build on this routing. Connection protocols (3.7) decide how these retrieval tools are exposed, for example as one shared MCP server. Progressive discovery (3.8) applies the same load-on-demand idea to tools and instructions. Evaluation datasets (4.2) turn Solveig's 120 questions into a proper test set, and diagnosing system issues (4.4) teaches you to suspect the retrieval step first when answers go wrong after an index refresh.
Key takeaways
- ✓ Choose retrieval per kind of question from two inputs, the shape of the data and the pattern of the query; one index for everything fails silently on codes, counts and chains of facts.
- ✓ Dense search matches meaning and misses exact identifiers, keyword search (BM25) matches exact tokens and misses paraphrase, and hybrid search with re-ranking combines them.
- ✓ Similarity search returns a top-k sample, so aggregation runs as a structured query through a tool and Claude explains the result.
- ✓ Hard constraints such as model, firmware, date or current version are metadata filters in the retrieval layer, not instructions asking Claude to ignore passages.
- ✓ Multi-hop questions need searches in sequence: agentic search lets Claude issue and refine them at the cost of turns and latency, and query rewriting maps the user's words onto the index.
- ✓ Retrieve wide, re-rank and send few, and skip retrieval when a small, stable corpus fits a cached prompt or the question needs no company data.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAR-P questions on Domain 3, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Integration alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 3 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.