Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 6 › Lesson 6.1

CCDV-F · Domain 6 · 11.0% of the exam · Lesson 6.1 · 22 min read

Context engineering: keep the window small and on task

How to keep an agent's context window useful: measure it, prune tool output, clear and compact old history, and isolate noisy work in subagents.

Written against skill 6.1 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

6.1.1 Why a data agent forgets the question it was asked

On Monday morning Keiko, head of e-commerce at a mid-sized online store for bike parts and cycling clothing, asks the team's new analysis agent a plain question: "Why did the return rate on our rain jacket double in September?" Sanna, the back-end developer who built the agent, gave it one tool, run_sql, which runs a read-only query against the orders database and returns the result. The agent starts well. It checks returns by size, then by colour, then by carrier, and reads the customers' free-text return reasons.

Forty tool calls later the answer arrives, and it is poor. It explains a fall in jacket SALES, which nobody asked about. It quotes a return rate from the seventh query, which double-counted exchanges and was corrected at call 19. Along the way it ran the same carrier query three times. Sanna's logs show why. Every query returned every row, and by call 40 each request carried about 300,000 tokens, the small chunks of text a model reads and bills by. Almost all of them were raw rows; Keiko's question was one short message near the top.

Nothing inside the model broke. The Claude API (the application programming interface your code calls) keeps nothing between requests. So your agent loop (the code that calls Claude, runs the tool Claude asks for, and calls again) resends the whole conversation every time. Everything the agent has seen stays in front of it, competing for attention. That working space is the context window: all the text Claude can take into account in one request, including the reply it writes. Context engineering is deciding what enters that window at each step, so that it stays small, relevant and pointed at the goal.

The same conversation at call 3 and at call 40

Call 3 about 12,000 tokens

System prompt and tools
Keiko's question
Two query results

Call 40 about 300,000 tokens

System prompt and tools
Keiko's questionone message near the top
39 query resultsmostly raw rows
A corrected figureand the stale one it replaced
Every call resends the whole history, so raw query results pile up until the question is one short message under hundreds of thousands of tokens of rows.

6.1.2 What fills the window, and why a bigger one does not help

Here is the belief that sends many teams the wrong way: if the agent struggles, give it more room. Many current Claude models accept up to a million tokens, so Sanna's 300,000 fit, and no error ever fired. The agent failed anyway. Size was never the problem; what filled the space was.

Everything in a request counts toward the window: the system prompt, the tool definitions, every message, and every tool result, image and document. The reply Claude writes counts too, including any thinking. And each tool result is counted again on every later call, because your loop resends it. Anthropic's documentation names what happens next context rot: as the token count grows, accuracy and recall degrade. Think of a detective's desk. Every printout from every lead stays on it, and the note with the client's actual question slides under the pile.

Two failures grow out of that pile before any limit is hit, and the exam guide names both. Context bloat is tokens that no longer help, such as raw rows already read, verbose logs and superseded results, which make every call slower and dearer. Context drift is the agent losing its goal, or leaning on stale material: it answers a neighbouring question, or reuses the figure that call 19 corrected.

Failure What you see Where it comes from
Bloat Input tokens in usage climb every call; each call is slower and costs more Raw tool output and old results kept in the history
Drift Answers wander from the question; stale or superseded facts reappear The goal is one early message among thousands of rows
Overflow A 400 "prompt is too long" error, or stop_reason: "model_context_window_exceeded" The input, or input plus reply, outgrows the window

Learn to recognise the first two by their symptoms, because a scenario often describes what the team sees rather than naming the failure. Overflow is the late, loud version of the same problem.

Managing the window starts with measuring it. Every response reports what the request consumed in its usage field. With prompt caching, the input is split across input_tokens, cache_read_input_tokens and cache_creation_input_tokens, and all three count toward the window. To check a request before you send it, the token counting endpoint (/v1/messages/count_tokens) returns its size. Sanna now logs the total on every call and sets a working budget well below the model's limit.

6.1.3 Prune tool output before it enters the window

Where did Sanna's 300,000 tokens come from? From run_sql, which returned every row it found. A query for returns by size and colour produced 2,400 rows. Claude needed the totals and a dozen rows that stood out, and paid for the rest on every later call.

The fix is tool output pruning: your tool code decides what a result contains before it reaches Claude. Anthropic's context-engineering guidance asks for tools that return token-efficient information, and your handler is where that is enforced. Return only the columns the question needs, cap rows with a default limit, and say how many were left out. Nudge Claude to aggregate in SQL instead of reading rows. And keep the full result outside the window, behind a reference Claude can use to fetch more.

That last part makes pruning safe: nothing is thrown away, only kept out of the conversation until needed. Anthropic's guide describes Claude Code analysing large databases the same "just in time" way: targeted queries, stored results, and only the slice it needs loaded into context. Here is Sanna's new handler. Look at the KEEP and PRUNE lines and the result_id: the full result goes to a store, only a preview goes back, and the id lets Claude ask for more.

MAX_ROWS = 20

def run_sql(query: str) -> str:
    rows, columns = db.execute(query)              # your read-only database client
    result_id = results_store.save(rows, columns)  # KEEP the full result outside the context
    preview = rows[:MAX_ROWS]                      # PRUNE: only what the next step needs
    return json.dumps({
        "result_id": result_id,                    # a reference for fetch_rows(result_id, offset)
        "row_count": len(rows),
        "columns": columns,
        "rows": preview,
        "note": f"Showing {len(preview)} of {len(rows)} rows. "
                "Aggregate in SQL, or call fetch_rows for more.",
    }, default=str)

The result_id does a second job, too. When Claude cites each finding with the id of its query, the figure from call 7 and its correction from call 19 stop being two anonymous numbers, and anyone can see which is current.

6.1.4 Clear what Claude has used, and keep notes outside

Pruning shrinks each result, but forty small results still pile up. Once Claude has read a result and drawn its conclusion, the rows have done their job, so why resend them on call 38? Anthropic calls clearing old tool results one of the safest, lightest-touch ways to shrink a context.

The API can do this for you with context editing, a beta feature. You add the clear_tool_uses_20250919 strategy to context_management.edits, with the beta header context-management-2025-06-27. Once the prompt passes a trigger (100,000 input tokens by default, or a number of tool uses), the API clears the oldest tool results first. It replaces each one with placeholder text that tells Claude it was removed, and keeps the most recent ones (3 by default). This happens on the server before the prompt reaches Claude; your code keeps the full history and has nothing to sync.

What if Claude later needs a figure from a cleared result? The answer is memory outside the window: notes kept in storage you control, which Claude reads back only when it needs them. Anthropic's memory tool gives Claude a directory of files under /memories to create, read and edit. It runs client-side: Claude requests each file operation, and your code carries it out against your own storage. Paired with context editing, Claude gets an automatic warning as the context nears the clearing threshold, so it can write down what matters before the rows go. Sanna's agent keeps one small file: the question word for word, the definition of return rate, each finding with its result_id, and the open questions.

Here is the request. Look at the trigger and keep lines, which decide when clearing starts and what survives it, and at the memory tool in tools.

response = client.beta.messages.create(
    model=MODEL, max_tokens=4096,
    system=SYSTEM,                                   # the question and definitions, on every call
    messages=messages,
    tools=[run_sql_tool, fetch_rows_tool,
           {"type": "memory_20250818", "name": "memory"}],  # notes that outlive clearing
    betas=["context-management-2025-06-27"],
    context_management={"edits": [{
        "type": "clear_tool_uses_20250919",
        "trigger": {"type": "input_tokens", "value": 60000},  # start clearing here
        "keep": {"type": "tool_uses", "value": 4},            # the 4 newest results stay
        "clear_at_least": {"type": "input_tokens", "value": 10000},
    }]},
)

Three options are worth knowing. exclude_tools lists tools whose results must never be cleared. clear_tool_inputs (false by default) also removes the tool calls, so leaving it off lets Claude still see which queries it ran. And clear_at_least matters with prompt caching: each clearing invalidates the cached prompt prefix, so make each one remove enough to be worth the cache rewrite. The response's context_management.applied_edits field reports what was cleared.

6.1.5 Compaction: replace the history with a summary

Clearing removes tool results but keeps every message, so a session that runs for hours still grows with Claude's reasoning and Keiko's follow-ups. At some point the conversation itself has to shrink. Compaction replaces the older part of the conversation with a summary Claude writes, then carries on from that summary. It is what a good colleague does before handing over a case: two pages of where things stand, not the whole folder.

On the API, compaction comes in two forms, both in beta on recent models. With on-demand compaction, which Anthropic's docs recommend wherever it is available, your code decides when. It sends the conversation with a compaction parameter, gets back one compaction block holding the summary, and puts that block first in messages in place of the turns it summarises. With threshold compaction, you add a compact_20260112 edit to context_management.edits instead. The API then writes the summary inside an ordinary request once input tokens reach your trigger. You append the response as usual, and the API ignores everything before the block.

Either way, the model you called writes the summary on Anthropic's servers, as an extra call you pay for. (If you build on the Claude Agent SDK, it compacts for you as the window nears its limit.) The art of compaction is choosing what to keep. The default prompt asks for whatever is needed to continue the task; custom instructions replace it completely, so they must name everything. Sanna's instructions also forbid tool calls: when tools are defined, the summariser occasionally tries to call one instead of writing, and then no summary comes back. Look at how the text pins the goal, the definitions, the source of each figure and the superseded numbers.

Summarise this analysis so it can continue in a fresh context. Keep, word for word: the user's question and the definition of return rate. Keep every finding with its figure and the result_id of the query behind it. Mark any figure that a later query corrected as superseded, with the corrected value. List the hypotheses already ruled out and the open questions. Leave out raw rows. Do not call any tools.

A summary can still drop an instruction that lived only in an early message. So keep standing rules, such as "cite a result_id for every figure", in the system prompt, which your code sends on every request and compaction never replaces. Here are the three ways of shrinking a history side by side.

Technique What it does to the history The risk to manage
Pruning, in your tool Keeps raw rows out from the start; the full result stays behind a reference Detail you did not return is invisible, so make it fetchable
Tool result clearing Replaces old tool results with placeholders past a trigger; messages stay A figure Claude never wrote down is gone
Compaction Replaces older turns with a summary and continues from it Anything the summary prompt did not ask for, including early instructions

Learn the risk column: when two options both shrink the context, the better one manages its risk.

6.1.6 Isolate noisy work in its own context

Even with pruning, clearing and compaction, one conversation that does everything stays crowded. Testing one hypothesis ("did the new supplier's batch leak?") can take a dozen queries, and three hypotheses make three dozen. The main conversation needs none of that search, only the answer to each question it asked.

Context isolation runs the noisy work in a separate context and brings back only the result. The first form is subagents. The main agent keeps the question, the plan and the findings, and starts one subagent per hypothesis. Each subagent begins with a fresh conversation that holds none of the parent's turns, runs its queries in its own context, and returns only its final message to the parent, as a tool result.

The Claude Agent SDK provides subagents ready-made. On the raw API you build one yourself: a tool, say investigate, whose handler runs a second agent loop with a fresh message list and returns only that loop's final answer. Anthropic's guidance describes subagents that spend tens of thousands of tokens exploring and hand back a condensed summary, often 1,000 to 2,000 tokens. Because a subagent sees nothing of the parent's history, its task must carry the question, the definitions, the date range and the output format.

A coordinator with three isolated investigations

Main agentthe question, the plan, the findings
Sizing subagent14 queries, returns 1 finding
Carrier subagent9 queries, returns 1 finding
Supplier subagent11 queries, returns 1 finding
Each subagent spends its own context on a dozen queries and hands back one short finding, so the main conversation holds the question and three findings instead of forty results.

The second form is a multi-step agentic workflow, where your code chains separate calls. Step one turns Keiko's question into hypotheses. Step two runs one investigation per hypothesis, each in its own request or loop with only its own inputs. Step three writes the report from the findings alone, never seeing a raw row. Choose the workflow when the steps are predictable and subagents when the agent must decide what to investigate next.

Isolation also works between tasks. On Tuesday Keiko asks an unrelated question about shipping costs. Continuing Monday's conversation would carry Monday's jacket results into Tuesday's answer. Start a fresh context for each new task, with at most a short handoff of what still matters.

6.1.7 The exam traps

Every trap here tries to fix a crowded context without changing what is in it.

  • ✗ Moving to a model with a larger context window to stop drift. ✓ Control what enters the window. Recall degrades as tokens grow, so more room only makes space for more noise.
  • ✗ Returning full query results and trusting the model to skip what it does not need. ✓ Prune in the tool: the needed fields and rows, a count of what was left out, and a reference to the full result.
  • ✗ Turning on prompt caching or batch processing to fix a bloated context. ✓ Reduce the tokens. Caching changes what you pay for tokens, not whether they occupy the window or distract the model.
  • ✗ Adding "always remember the original question" to the prompt, or replaying earlier transcripts. ✓ Pin the goal in the system prompt, and clear, compact or isolate the material that buries it.
  • ✗ Compacting with a vague summary prompt, or saving summaries as trusted memory. ✓ Say what the summary must keep (goal, decisions, sources, superseded figures), and treat text that came from tool output as data, never as an instruction.
  • ✗ Running one long session across unrelated questions or customers. ✓ Start a fresh context per task with a short handoff. Otherwise stale results from the last task leak into the next.

Four tempting fixes, one real one

A bigger context windowroom for more noise
A bigger modelthe same buried question
Prompt cachingcheaper tokens, same tokens
"Remember the question"one more line in the pile
Control what enters the windowprune, clear, compact, isolate
More room, a stronger model, cheaper tokens and a stern reminder all leave the same noise in the window; only changing what enters it fixes drift and bloat.

6.1.8 Put it together: fix the forgetful analyst

You now have every piece: measure the window, prune at the source, clear what has been used, keep notes outside, compact the conversation and isolate noisy work. The quickest way to make it stick is to watch a small analysis agent drift, fix it, then break the fix.

The rest of this domain works on the two ends of the same request. Prompt engineering (6.2) covers how to word the standing instructions you now keep in the system prompt, and where each kind of instruction belongs. Output handling (6.3) covers checking the agent's report before anyone acts on it. And cost and token management (5.4) turns the token counts you logged into a budget.

Key takeaways

  • ✓ Everything a request carries counts toward the context window and recall degrades as it fills, so measure usage on every call and manage the window from the first turn.
  • ✓ Bloat is tokens that no longer help and drift is the agent losing its goal or reusing stale results; a bigger window or model fixes neither.
  • ✓ Prune tool output in your tool code: return the rows and fields the next step needs, say what was left out, and keep the full result behind a reference.
  • ✓ Context editing clears old tool results on the server past a trigger; pair it with memory outside the window so the findings Claude needs survive.
  • ✓ Compaction replaces older turns with a summary, so tell it what to keep and put standing rules in the system prompt, which it never replaces.
  • ✓ Isolate noisy work in subagents or separate workflow steps that return condensed findings, and start a fresh context for each unrelated task.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

18 CCDV-F questions on Domain 6, free

Every question in the bank is tagged to a domain, so you can drill 18 questions on Prompt and Context Engineering alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 6 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources