Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 2 › Lesson 2.4

CCAR-P · Domain 2 · 13% of the exam · Lesson 2.4 · 21 min read

Optimising the context window: select, compact and measure tokens

Why a bigger window is not the fix, how to choose what enters each request, keep long sessions lean with compaction and clearing, and measure tokens.

Written against objective 2.4 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.4.1 Why the assistant got worse as the shift went on

Larkspur Aero Maintenance services regional turboprops for three airlines. When a fault will not clear, a technician on the hangar floor opens an assistant on a tablet and troubleshoots with it for hours. The rhythm never changes: describe the symptom, get the next isolation step, run the test, report the reading, ask again. Behind the assistant sit the maintenance manuals for each aircraft type, thousands of pages of procedures, limits, wiring diagrams and service bulletins. Emeka, the solution architect Larkspur brought in, inherited a pilot that technicians trusted in the first hour of a session and doubted by the third.

The pilot ran on a model with a 200K-token window (tokens are the word pieces a model reads, writes and bills by) and did the generous thing on every question. It sent each manual chapter that looked relevant in full, resent the whole session transcript, and kept every raw tool output: log queries returning two hundred past write-ups, sensor dumps, parts lookups. Three hours into chasing an intermittent low-pressure caution in hydraulic system 2, the assistant quoted a pressure limit from a different variant's procedure. It asked Thabo, the lead technician, to repeat a test he had run at the start. Answers took most of a minute to begin, and in hour four requests failed with "prompt is too long".

The vendor proposed a model with a bigger window. When Emeka replayed recorded sessions on it, the failures stopped, the wrong-variant answers did not, and the longest sessions cost several times more. The real problem was never the size of the window; it was what filled it. The one procedure that mattered competed with dozens that did not, and the reading from hour one sat buried under a hundred thousand tokens of later chapters and log output. What Larkspur needed was context engineering: deciding, request by request, which tokens the model sees, keeping long sessions lean, and measuring the spend.

The same question, two ways to fill the window

The pilot

Whole chaptersboth aircraft variants
Full transcriptevery turn, word for word
Every raw tool outputkept all session
Slow, costly, wrong variant

The redesign

Sections for this faultthis aircraft's variant only
Case summaryplus the last few turns
Newest tool resultsolder ones cleared
Fast, cheaper, cites the right task
The pilot sent everything that might help; the redesign sends what this question needs and keeps the rest outside the window until it is asked for.

2.4.2 One budget, shared by input and output

Here is the question that catches experienced engineers out: what actually counts against the window? Most people think of the question and the documents. In fact everything in the request counts, from the system prompt to every tool result, and so does the reply. The context window is all the text the model can reference while it generates a response, INCLUDING that response, and its size is fixed per model.

Think of a workbench on the hangar floor. Its size is fixed, and the parts for this job must share it with whatever the last job left, with room to assemble the result. A technician who never clears the bench does not run out of space all at once. They spend longer finding the right part, and now and then they pick up the wrong one.

What fills the window What makes it grow The architect's lever
System prompt Rules, examples and policy text added over time Keep it minimal but complete; reuse the stable part
Tool definitions Every attached tool, resent on every turn Attach only the tools the role needs; defer rarely used ones
Conversation history Every turn, resent with each request A structured case summary plus recent turns; compaction
Retrieved documents Whole chapters instead of relevant sections Select per question, filtered to the case
Tool results Raw logs and query dumps kept all session Compact results; clear old ones
Thinking and the reply Task difficulty, effort, max_tokens Reserve enough output for thinking plus the answer

Memorise the six rows; a context problem in production usually hides in one of them. The max_tokens parameter caps what the model may generate this turn, and thinking tokens (the reasoning Claude writes before it answers) are part of that cap and billed as output. On current Opus, Sonnet and Fable models the API keeps earlier thinking blocks by default, so reasoning from earlier turns becomes input for later ones.

The limits behave in two ways. If the input alone exceeds the window, the request fails with a 400 "prompt is too long" error. On Claude 4.5 and newer models, a request whose input plus max_tokens exceeds the window is still accepted. If generation then reaches the limit, it stops with stop_reason (the response field that says why generation ended) set to "model_context_window_exceeded". Emeka broke down one hour-three request: about 6,000 tokens of system prompt and tools, 70,000 of manual chapters, 45,000 of transcript and 60,000 of tool output. That is some 181,000 tokens against a 200K window, and only a few thousand of them were evidence the question needed.

2.4.3 Why a bigger window is not the fix

It is tempting to believe that if something fits, including it can only help. Resist it. As the token count grows, accuracy and recall degrade, an effect Anthropic's documentation calls context rot. The cause is structural. Every token attends to every other token, so n tokens create n² relationships, and the model's attention is stretched thinner as the context grows. Anthropic's engineers describe an attention budget that every new token draws down. The result is a gradient rather than a cliff: the model stays capable at long lengths but becomes less precise at picking one fact out of many.

Larkspur's wrong-variant answer shows what that loss of precision looks like. Two variants of the aircraft have fault-isolation procedures that differ in one pressure limit and one connector number. With both chapters in context, the right procedure and its near-twin competed, and the model sometimes quoted the wrong one. A bigger window would have kept both in place.

So when does a long-context model win? Current Opus, Sonnet and Fable models have a 1M-token window, while Claude Haiku 4.5 has 200K. The 1M window is billed at the standard per-token rate, which sounds reassuring until you do the arithmetic. A 900,000-token request costs a hundred times as much input as a 9,000-token one. Every token is also processed before the first word of the answer, so latency grows with it.

A large window earns its cost when a task truly needs many documents at once, such as reconciling one aircraft's modification history against a new service bulletin. It could never be Larkspur's whole answer, because the manual set for one aircraft type runs to several million tokens. Emeka did move the assistant to Claude Sonnet 5.5, whose window is 1M, but as headroom for those rare cross-document questions; everyday requests still have to be small.

Why "it fits, so send it" fails

Irrelevant pages competeaccuracy and recall fall
Every token is billed900K of input costs 100 times 9K
Every token is processedthe first word arrives later
The whole manual still will not fitmillions of tokens
Curate first, size the window seconda big window is headroom, not focus
A bigger window removes the error message, not the costs of an overfull request; curating what enters removes all of them.

2.4.4 Choosing what enters: select per question, documents first

If the whole manual cannot go in, the real design question is which of its pages THIS question needs: usually a handful of sections. Emeka replaced "send what might help" with a context assembly step that builds each request from parts, in the same order every time.

Assembling one troubleshooting request

SELECTthe manual sections for this fault
FILTERthis aircraft's variant; drop near-duplicates
SUMMARISEcase state plus the last few turns
ORDERdocuments first, question last
RESERVEmax_tokens for thinking and the answer
Each request is built for one question: the right sections for this aircraft, the case so far in summary, documents above the question, and room reserved for the answer.

Selection and filtering remove the near-twin procedure before the model ever sees it. Whatever retrieval sits behind SELECT, it must return sections rather than chapters, and FILTER must apply the aircraft's variant. The case summary, which the application updates after each recorded test, replaces the raw transcript with what one technician would hand the next at a shift change.

Order matters more than most teams expect. For inputs of 20,000 tokens or more, Anthropic's guidance is to put long documents at the top and the query at the end. In its tests, that ordering improved response quality by up to 30 percent, especially with complex multi-document inputs. The same guidance says to wrap each document in tags with its source, and to ask for the relevant quotes before the answer. The assembled request looks like this; look at the source line on each document, and at the case summary and question, which come last.

<documents><document index="1"><source>Fault isolation manual, task 29-10-05, revision 42, variant -300</source><document_content>...</document_content></document><document index="2"><source>Maintenance manual, task 29-11-01, revision 42</source><document_content>...</document_content></document></documents>
<case_summary>Tail LA-214, variant -300. Fault: HYD 2 LO PRESS caution, intermittent, in climb. Done: transducer connector inspected and cleaned (10:40); pump case drain flow within limits (11:55). Open: caution returned at 13:10; transducer replacement not yet done.</case_summary>
Technician's question: the caution came back after the connector was cleaned. What is the next isolation step, and which limit applies? First quote the lines from the documents you rely on, then answer, citing the task number and revision.

2.4.5 Keeping a long session lean

Selection controls what you ADD each turn, but a session also accumulates: turns, tool results, thinking. Four mechanisms each remove a different kind of weight, and the architect's job is to match them to what is actually growing.

Compaction replaces older turns with a summary Claude writes, and the conversation continues from it. It is in beta on Claude 4.6 and later models, and the docs name it the main strategy for long conversations. Your code can request it on demand, keeping the last few turns word for word, or let the API run it at a token threshold. What the summary keeps is the design decision, so Larkspur replaces the default prompt with its own. Its instructions say: keep every reading with its time, every test and its result, every part changed, and the current hypothesis. Anthropic advises tuning them for recall first and precision second: a detail dropped in hour two may matter in hour five.

Tool result clearing, a strategy of context editing (also in beta), removes the oldest tool results once the input passes a threshold you set. It keeps the most recent few and leaves a placeholder where each cleared result was, so the model still knows it made the call. Clearing runs on the server, so your application keeps the full history. A second strategy, thinking block clearing, does the same for earlier reasoning carried from turn to turn.

The memory tool lets Claude save notes to files under /memories: Claude requests the file operations, and your application executes them against storage you control. Notes survive clearing, compaction and even a new session, which at Larkspur means a shift change. With memory and clearing both enabled, Claude is warned as the context nears the clearing threshold, so it can save findings first.

Subagents, separate Claude calls to which the main agent hands one focused task, keep exploration out of the main window. Searching eighteen months of fleet write-ups for the same caution might read tens of thousands of tokens; a subagent does it in its own context and returns a condensed summary, often 1,000 to 2,000 tokens. Emeka's brief asks it to return the matching write-up numbers too, because the main agent can cite only what comes back.

Technique When it wins What it costs
Compaction Long back-and-forth sessions whose thread must continue An extra model call each time; whatever the summary omits is gone
Tool result clearing Agent loops where old results are dead weight A cleared result must be fetched again if needed; each clearing invalidates the cached prefix
Memory tool Facts that must outlive clearing, compaction or the session Storage you build and secure; extra tool turns
Subagent Exploration that reads a lot to return a little Extra calls and latency; the main agent sees only the summary

Emeka's request configuration puts clearing and memory side by side. Look at trigger, keep and exclude_tools, and at the memory tool in the tool list.

response = client.beta.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=8000,                      # output reserve: thinking plus the answer
    betas=["context-management-2025-06-27"],
    system=SYSTEM_PROMPT,
    tools=[SEARCH_MANUAL, GET_FAULT_HISTORY, GET_TASK_CARD,
           {"type": "memory_20250818", "name": "memory"}],   # findings survive clearing
    messages=history,
    context_management={"edits": [{
        "type": "clear_tool_uses_20250919",
        "trigger": {"type": "input_tokens", "value": 80000},  # start clearing here
        "keep": {"type": "tool_uses", "value": 4},            # newest four results stay
        "clear_at_least": {"type": "input_tokens", "value": 15000},
        "exclude_tools": ["get_task_card"],                   # the open task card never goes
    }]},
)

2.4.6 Measure the budget before you defend it

Larkspur's operations director asked Emeka a fair question: how do you know where the tokens go, and what will the redesign cost per session? An architect answers with measurements, not estimates from word counts.

Before sending, the token counting endpoint takes the same inputs as a message (system prompt, tools, images, PDFs) and returns the input token count. It is free, rate-limited separately from message creation, and returns an estimate that can differ slightly from the real request. Count against the model you will use: Claude 4.7 and later models use a newer tokenizer that turns the same text into roughly 30 percent more tokens than earlier models. Larkspur's preflight counts each request and, when it is over budget, requests an on-demand compaction before sending, not after a failure.

After the response, every reply reports a usage field: input_tokens and output_tokens, plus cache_read_input_tokens and cache_creation_input_tokens when caching is on, and all three input fields count toward the window. With clearing enabled, context_management.applied_edits reports how many tool uses and input tokens were cleared. Compaction is billed too: its summarisation call appears in usage.iterations, so sum that array, not just the top-level fields. A reply that reaches max_tokens ends with stop_reason: "max_tokens", cut off mid-answer, so log it rather than ignore it. Emeka wrote the budget down as an artefact the director could sign.

Token budget, hangar assistant, one troubleshooting request (Claude Sonnet 5.5, 1M window).
Input target: 60,000 tokens at the 95th percentile (p95); alert at 90,000; preflight compacts above 120,000.
System prompt and tool definitions: about 7,000, stable across the session.
Case summary up to 2,000; last six turns word for word.
Manual sections up to 30,000: at most eight, filtered to the aircraft's variant.
Tool results: newest four kept; older ones cleared above 80,000 input tokens; the open task card is never cleared.
Output reserve: max_tokens 8,000 for thinking plus the answer; every stop_reason "max_tokens" is logged and reviewed.
Evidence: usage fields and applied_edits logged per request; p95 input tokens reported weekly against target.

2.4.7 The exam traps

Every trap below either adds tokens that do not help or removes tokens that do.

  • ✗ Moving to a bigger window or a bigger model when long sessions degrade. ✓ Curate what enters. The usual cause is a crowded window, not a small one; a larger window keeps the clutter and raises cost and latency.
  • ✗ Truncating the oldest turns or each document at a fixed token count. ✓ Select by relevance and summarise with instructions that name what must survive. Blind truncation drops the reading taken in hour one.
  • ✗ Loading whole chapters "to be safe". ✓ Send the sections this question needs, filtered to the case. Near-duplicate procedures compete with the right one.
  • ✗ Keeping every raw tool result for the whole session. ✓ Return compact results, clear stale ones, exclude those that must stay, and save findings to memory first.
  • ✗ Putting the question first and the documents after it. ✓ Put long documents at the top and the question last, with each document tagged by source; Anthropic's tests show better answers that way.
  • ✗ Budgeting from word counts or counts taken on another model. ✓ Count tokens against the target model and log the usage fields of every response, because tokenizers differ between model generations.

2.4.8 Put it together: budget a long troubleshooting session

You now have every piece, and the fastest way to make them stick is to watch a crowded context fail on questions a curated one answers.

Prompt reuse (2.5) makes the stable part of each request cheaper with caching. Retrieval pipeline design (3.5) and retrieval strategies (3.6) decide whether SELECT finds the right sections. Progressive discovery (3.8) weighs loading capabilities on demand against loading them up front, and cost-performance optimisation (4.5) turns your token budget into the numbers finance tracks.

Key takeaways

  • ✓ The context window is one fixed budget for the system prompt, tools, history, documents, tool results, thinking and the reply; max_tokens reserves the output share.
  • ✓ More context is not free quality: accuracy and recall degrade as tokens grow, and a bigger window adds cost and latency without removing the clutter.
  • ✓ Assemble each request for one question: select the relevant sections, filter to the case, summarise the session, and put long documents above the question.
  • ✓ Keep long sessions lean with compaction for old turns, tool result clearing for stale output, the memory tool for durable facts, and subagents for exploration.
  • ✓ Every technique loses something, so state what it drops and protect what the task needs, such as readings, tests done and the open task card.
  • ✓ Measure before you optimise: count tokens against the target model, log the usage fields and applied edits, and treat stop_reason: "max_tokens" as a truncated answer.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

24 CCAR-P questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 24 questions on Claude Models, Prompting & Context Engineering alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources