Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 3 › Lesson 3.5

CCAR-P · Domain 3 · 19% of the exam · Lesson 3.5 · 22 min read

Designing a RAG pipeline: chunking, metadata and indexing

How to design a RAG pipeline: chunk by document structure, label every chunk, index for meaning and exact terms, keep the index current, cite sources.

Written against objective 3.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

3.5.1 Why the caseworker assistant cannot just read the manuals

A caseworker at the Tarrow Revenue Authority, a national tax agency, has a taxpayer on the phone. The taxpayer worked from home for part of 2025 and wants to know whether a share of the heating bill is deductible. The answer sits somewhere in 30,000 pages of guidance manuals, published rulings and form instructions, updated every month. The relevant manual section was revised in July, so the answer for 2025 differs from the answer for 2026. An experienced caseworker knows where to look; a new one searches for twenty minutes and may still quote the wrong year.

The Authority has engaged Rhiannon, a solution architect, to design a Claude assistant for its caseworkers. A plain model call cannot answer this well. Claude knows how home-working deductions tend to work in general, but it has never seen Tarrow's manuals, so without them it offers a plausible guess about someone else's tax law. Nor can Rhiannon paste the manuals into the prompt. The context window, the most text Claude can read in one request, is at most 1M tokens on current models, roughly 555,000 words; 30,000 pages is many times that.

The pattern for this is retrieval-augmented generation (RAG): at question time, your application finds the few passages that answer the question and hands them to Claude with it. A small knowledge base can skip RAG and go into the prompt whole; Anthropic puts that line at about 200,000 tokens, some 500 pages. Tarrow's corpus is sixty times that. So Rhiannon's real work happens before any question is asked. She decides how documents are cut into passages (chunking), what labels each passage carries, how passages are made searchable (indexing), and how next month's revision reaches the index.

Answering from memory versus answering from the guidance

Plain model call

The caseworker's question
Claude answers from general knowledge
Plausible, uncited, possibly the wrong year

Retrieval first

The caseworker's question
Your application retrieves matching sectionsfiltered to tax year 2025
Claude answers from them, with citations
Without retrieval Claude can only offer general knowledge; with it, Claude answers from the Authority's own sections for the right tax year and points to them.

3.5.2 The pipeline: two paths and a way back in

Teams new to RAG often picture one step: search the documents, then ask Claude. That hides where quality is decided. A RAG system is two paths that meet in the index, and an architect designs both.

The ingestion path runs offline, once per document version. It pulls each document from its source of record and parses it, keeping headings, sections and tables. It cuts the text into labelled chunks, turns each chunk into an embedding (a list of numbers that encodes its meaning) and writes it to the indexes. The query path runs online, once per question. It retrieves candidates, re-ranks them with a model that scores each one against the question, and has Claude answer from the best few, with citations.

The two paths of a RAG pipeline

Ingestion path once per document version

Ingest and parsekeep headings, sections, tables
Chunk and labelmetadata, a context sentence
Embed and indexfor meaning and for exact terms

re-run for every changed document

Query path once per question

Retrievefiltered, from both indexes
Re-rankkeep the best few
Generatea cited answer, or "not covered"
The query path can only choose among the chunks the ingestion path produced, so most quality decisions are made on the left.

Think of a library. The librarian at the desk can be brilliant, but if the cataloguer tore chapters in half and left the edition off the spine, the best librarian still hands you the wrong book. The query path is the librarian; ingestion is the catalogue. A table split in half at ingestion cannot be rejoined at query time, and a chunk with no effective date cannot be filtered by one.

The loop at the bottom of the left column is the re-index path, and it is part of the design, not an operations afterthought. Each monthly publication event triggers ingestion for the changed documents only, writing to every index in one step so they never disagree. Old versions are end-dated, not deleted: a caseworker settling a 2024 return needs the 2024 guidance. A withdrawn ruling, by contrast, must stop being retrievable at once. A nightly check compares document versions in the source with those in the index and raises an alarm on any gap.

3.5.3 Chunking: cut where the meaning breaks

It is tempting to pick one chunk size, say 500 tokens, and split everything by count. On Tarrow's guidance that fails at once. The home-working manual has a rate table with one column per tax year, and a count-based cut can put the header row in one chunk and the figures in the next. The retrieved chunk then holds the right numbers and no way to tell which year they belong to.

There are four ways to cut, and each wins for a different kind of content.

Chunking option When it wins What it costs
Fixed size with overlap (each chunk repeats the end of the one before) Uniform prose with no reliable structure; a quick first version Cuts through sentences, sections and tables; the overlap duplicates text in the index and in the context
Structure-aware (by section and heading, tables kept whole) Documents with headings and numbered sections, such as manuals A parser per format; sections vary in size, so long ones are split again and short ones merged
Per record Self-contained units: one ruling, one form line, one FAQ entry Only works where the source has records; an oversized record still needs splitting inside it
Semantic (split where the topic shifts) Long unstructured text, such as transcripts or old memos Extra compute at ingestion; boundaries are less predictable and harder to cite

Size is the second decision, and it trades precision against context. A small chunk holds one idea, so its embedding is sharp and a match precise. But it may lack what the answer needs, such as the exception in the next paragraph or the tax year in the heading. A large chunk carries that context, but its embedding averages several topics, so matches blur, and each chunk costs more tokens, so fewer fit. Anthropic's write-up describes typical chunks as no more than a few hundred tokens; what decides is how much text a typical question needs.

Rhiannon's design follows the content types. Manuals are cut by numbered subsection, each chunk carrying its heading path. Tables stay whole with their caption and header row; a table too large for one chunk is split by rows with the header repeated in every piece. Each ruling is one record, and each form instruction is cut per form line. Fixed-size chunks survive only for scanned legacy circulars with no usable structure.

3.5.4 Labels that travel with every chunk

Here is the question that trips people up. The 2025 and 2026 versions of the heating section share almost every word. When a caseworker asks about 2025, which one does vector search return? Whichever scores fractionally closer: close to a coin toss. Similarity measures meaning, and the two versions mean nearly the same thing. Which version is in force for the tax year is not a matter of meaning at all. It is a fact about the document, and facts belong in metadata: fields stored with each chunk that the index can filter on exactly.

Look at the two dates, which select the version in force, and at audience and access, which drop what the caller may not see.

{
  "chunk_id": "hwm-4.2.3@2026-07",
  "doc_id": "home-working-manual",
  "doc_type": "manual",
  "section_path": "4 Expenses > 4.2 Working from home > 4.2.3 Heating and power",
  "version": "2026-07",
  "effective_from": "2026-01-01",
  "effective_to": null,
  "audience": "caseworker",
  "access": "general",
  "source": "tra://home-working-manual/4.2.3@2026-07",
  "text": "Heating and power may be claimed in proportion to ..."
}

These fields do four jobs. They FILTER by date: superseded chunks stay in the index but reach Claude only for cases in their years. They FILTER by access: your retrieval layer applies the caller's entitlements, so a fraud-investigation procedure never enters a general caseworker's context. That only works if chunks are cut at access boundaries, so each label is true of the whole chunk. They make CITATIONS checkable: a section path and version point to something a caseworker can open and an auditor can reconstruct. And doc_id plus version tell the re-index path exactly which chunks a revision supersedes.

Filter before you rank, not after. Retrieve the top 20 by similarity, then discard ineligible chunks, and you may keep three, or none. Filtering inside the index means ranking only among chunks the caller may see and the case may use.

3.5.5 Indexing for meaning and for exact terms

Tarrow's corpus has to be findable in two ways. Caseworkers describe most of it in their own words: "can someone who works at the kitchen table claim heating?" shares almost no words with "4.2.3 Heating and power". But it is also full of identifiers, such as form TR-114, ruling numbers and section references, where only the exact string will do.

A vector index stores the embeddings and finds the chunks nearest in meaning. It handles paraphrase, but asked about TR-114 it can return forms on similar topics. A keyword index scored with BM25 (Best Matching 25), a long-established function that ranks chunks by the query's exact terms and weights rare ones more, handles identifiers. A hybrid builds both over the same chunks and fuses their ranked lists; Anthropic's tests found embeddings plus BM25 beat embeddings alone.

Chunks also lose something when they leave their document. "The allowance is capped at the amount in Table 4B" never says which allowance or which year. Contextual retrieval fixes this at ingestion. Your pipeline sends Claude the whole document and one chunk and asks for a short context that situates the chunk, usually 50 to 100 tokens, then prepends it before embedding and before keyword indexing. The context is specific to the chunk: Anthropic saw very limited gains from attaching a generic document summary instead.

In Anthropic's tests, contextual embeddings cut the top-20 retrieval failure rate by 35%, adding contextual BM25 cut it by 49%, and adding a re-ranking step as well cut it by 67%. The price is a Claude call per chunk at ingestion. Prompt caching, which reuses an already-processed prompt prefix at a fraction of the input price, lets each document be read once for all its chunks, and offline ingestion can use the half-price Message Batches API. A revision can change any chunk's context, so the re-index path regenerates contexts for the whole document.

What one chunk goes through before it is searchable

CHUNK"capped at the amount in Table 4B"
CONTEXTClaude reads the whole manual, writes 50 to 100 tokens
PREPEND"Home-working manual, 2026 rules, heating allowance..." + chunk
INDEXembed for meaning, BM25 for exact terms
Contextual retrieval adds a short, document-aware sentence to each chunk, so both the vector index and the keyword index see the words the chunk itself leaves out.

The embeddings come from a separate provider, because Anthropic does not offer an embedding model. Its docs point to Voyage AI (general models, domain models for law and finance, contextualized chunk models, rerankers) and advise assessing several vendors on domain fit, speed at your scale and customization. Two rules hold whichever you pick. Embed documents and queries with the same model, marking which is which (Voyage's input_type parameter). And a change of embedding model means re-embedding the whole corpus, because vectors from different models generally do not share a space.

Index option When it wins What it costs
Vector only A prose corpus searched in paraphrase Misses exact identifiers: form numbers, ruling IDs, section references
Keyword (BM25) only A corpus searched mostly by identifier or exact term Misses paraphrase and synonyms
Hybrid, both over the same chunks A mixed corpus of prose and identifiers, like Tarrow's Two indexes to build, update and keep in step
+ contextual retrieval Chunks that lose their meaning outside their document A Claude call per chunk at ingestion, repeated for the whole document when it changes
Contextualized chunk embeddings (Voyage's voyage-context-4) Document-aware vectors without a Claude call per chunk Helps the vector index only; the keyword index still sees the bare chunk

3.5.6 Grounding the answer: citations and "the guidance does not say"

Now suppose retrieval works: the right sections reach Claude, yet the answer blends them with general tax knowledge, and nobody can tell which sentence came from where. A caseworker cannot give a taxpayer advice they cannot trace to guidance. Asking Claude in the prompt to "cite your sources" only produces citations the model writes itself, which nothing checks.

The Claude API has two built-in mechanisms. With citations enabled on document blocks, the answer's text blocks carry pointers to the exact passages they draw on. Search result content blocks are built for RAG. Each retrieved chunk becomes a search_result block with a source (any stable string, such as a URL or internal ID), a title and a content list of text blocks. With citations enabled, Claude cites them the way it cites web search results. The docs guarantee that these pointers are valid, and cited_text costs no output tokens. In Anthropic's evaluations, the feature also cited the most relevant quotes more often than prompting did.

Here is a sketch of the search tool behind Tarrow's assistant. Note three lines: the filter applied before ranking, the source built from the chunk's metadata, and the plain text block returned when nothing matches.

def search_guidance(query: str, case_date: str, caller: Caller) -> list[dict]:
    hits = hybrid_search(query, filters={
        "in_force_on": case_date,              # the version in force for this case
        "access": caller.entitlements,         # FILTER before ranking, not after
    })
    if not hits:                               # let Claude say so, not guess
        return [{"type": "text", "text": "No guidance found for this question."}]
    return [{
        "type": "search_result",
        "source": h.source,                    # e.g. tra://home-working-manual/4.2.3@2026-07
        "title": f"{h.doc_title}, {h.section_heading}",
        "content": [{"type": "text", "text": p} for p in h.paragraphs],  # finer citations
        "citations": {"enabled": True},        # off by default; same on every result
    } for h in rerank(query, hits)[:20]]

Claude cites whole text blocks, so splitting a chunk into paragraphs gives finer citations. The empty case follows the docs: return a plain text block, not an error, and Claude explains the empty result. Pair it with a system prompt that permits "the guidance I found does not cover this" and restricts Claude to the guidance provided, two techniques from Anthropic's hallucination guide. For Tarrow, "not covered, refer to the technical team" is a correct answer; an uncited paragraph from general knowledge is not.

3.5.7 The exam traps

Every trap here fixes a pipeline problem somewhere other than where it happens.

  • ✗ One fixed chunk size for the whole corpus because it is simple. ✓ Chunk by structure per content type, tables whole, one record per chunk. Count-based cuts separate numbers from the headers that give them meaning.
  • ✗ Trusting similarity to pick the right version, or a prompt telling Claude to ignore superseded or restricted passages. ✓ Store dates, versions and access labels as metadata and filter before ranking. Near-identical versions look alike to an embedding, and text already in the context cannot be unseen.
  • ✗ A vector index alone over a corpus full of form numbers and ruling IDs. ✓ Build a BM25 keyword index over the same chunks and fuse the two. Embeddings capture meaning, not exact strings.
  • ✗ More chunks or a bigger model to fix poor retrieval. ✓ Fix the chunks, labels and index. No model can answer from a passage it was never given.
  • ✗ Deleting the old version when a document is revised. ✓ End-date it and filter on the case's date; cases about earlier years need the guidance that applied then.
  • ✗ Asking Claude to list its sources in prose. ✓ Send search result blocks with citations enabled, and permit "not covered". API citations point to passages that were sent; self-reported sources are claims.

Four tempting fixes, one real one

A bigger modelsame passages, same gap
More chunksmore cost, more distraction
"Ignore old versions"the text is already in context
"Cite your sources"self-reported, unchecked
Fix the pipelinestructured chunks, metadata filters, hybrid index, API citations
A bigger model, more context, prompt instructions and self-reported sources all leave the pipeline defect in place; the fix belongs to the stage that failed.

3.5.8 Put it together: design Tarrow's ingestion and index

You now have every piece of the design. Here is how Rhiannon records it for the Authority's architecture board.

Decision: Guidance retrieval for the caseworker assistant (ingestion and index design).
Requirement: Answers must use the guidance in force for the case's tax year, respect access labels, and cite section and version; releases are monthly.
Chunking: manuals by numbered subsection with heading path; tables whole with header row, split by rows with header repeated; one chunk per ruling; one per form line; fixed size only for unstructured legacy circulars.
Metadata: doc_id, section_path, version, effective_from, effective_to, audience, access, source.
Index: hybrid vector and BM25, both over contextualized chunks; effective-date and access filters applied inside the index before ranking; a re-ranking step keeps the best 20 for Claude.
Sync: each publication event re-ingests the changed documents, contexts included, into both indexes at once; superseded versions end-dated, withdrawn rulings removed; nightly version reconciliation.
Evidence: on 300 questions from caseworkers, the answering section was in the top 20 for 96% with this design against 81% for fixed chunks and vector search.
Revisit when: a new content type is added, the embedding model changes, or top-20 recall drops below 92%.

Retrieval strategies (3.6) take this index and match the retrieval method to each query pattern, including when a database lookup beats any search. Progressive discovery (3.8) decides how much of what you can retrieve enters the context up front. And diagnosing failures (4.4) finds which stage of this pipeline broke when answers go wrong.

Key takeaways

  • ✓ Claude can only be as right as the passages retrieved, and the ingestion path (parse, chunk, label, embed, index) sets that ceiling before any question is asked.
  • ✓ Every document change must flow back through ingestion: re-index changed documents into all indexes at once, end-date superseded versions and remove withdrawn ones.
  • ✓ Chunk along the document's own structure, keep tables whole, use one record per chunk where records exist, and size chunks by what a typical question needs.
  • ✓ Store source, section, version, effective dates, audience and access on every chunk, filter on them before ranking, and build citations from them.
  • ✓ Build vector and BM25 indexes over the same chunks for a corpus of prose and identifiers; contextual retrieval raises accuracy at an ingestion cost your evals must justify.
  • ✓ Anthropic has no embedding model and its docs point to Voyage AI; embed queries and documents with one model, and re-embed the corpus if you change it.
  • ✓ Ground answers with search result blocks or documents with citations enabled, and let Claude say the documents do not answer.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

36 CCAR-P questions on Domain 3, free

Every question in the bank is tagged to a domain, so you can drill 36 questions on Integration alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 3 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources