Home › Study guides › CCDV-F › Domain 1 › Lesson 1.3
CCDV-F · Domain 1 · 14.7% of the exam · Lesson 1.3 · 21 min read
Agent patterns and the frameworks that package them
The patterns inside most agents (tool-use loops, subagents, memory, context-window management) and what frameworks like LangGraph package and hide.
Written against skill 1.3 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
1.3.1 Why one model call cannot work a case file
Picture a litigation associate handed a commercial lease dispute on Monday: about 400 documents, from the lease itself to emails, invoices and inspection reports. The partner wants one answer: when did the landlord first learn about the leaking roof? Nobody holds 400 documents in their head. The associate searches the file, hands two boxes to juniors, keeps notes, and picks the work up again on Thursday from those notes.
Ingrid and Tomasz, two developers at a small legal-tech company, are building a case-file research assistant to do this job with Claude. A single call to the Messages API, the endpoint your code calls to talk to Claude, cannot do it. The call cannot open the document store; Claude can only reply to what you send. Four hundred documents may not even fit in the context window, the most text Claude can take in at once, measured in tokens (the small chunks of text a model reads and bills by). Even material that does fit is recalled less reliably as the request grows. And each request is a blank slate, so Thursday knows nothing of Monday.
The answer is a small set of proven patterns. An agent is Claude using tools in a loop, and that tool-use loop is the engine. Three more patterns keep it working as the job grows: subagents that do the reading in their own context, memory that outlives a session, and context-window management that keeps every request small and relevant. Frameworks such as LangGraph, Strands and PydanticAI package these patterns, and the team must decide whether to use one.
1.3.2 The tool-use loop: the engine under every pattern
Here is the question to settle before any pattern or framework: what actually runs inside an agent? Always the same cycle. Your code sends Claude the conversation plus a list of tools. Claude replies with a final answer or with a structured request to use a tool. Claude never runs anything itself: your code runs the tool, appends the result to the conversation and sends everything again. The reply's stop_reason field says which case you are in, and "tool_use" means go round again.
For the case-file assistant, a run might look like this. Claude requests search_case_files for "roof leak", and your code returns twelve document ids with titles. Claude requests read_bundle on four of them, looks at the findings, and searches again for the landlord's replies. Nobody scripted that order: Claude chose each step from the previous result, which makes this an agent rather than a fixed pipeline.
The tool-use loop every agent runs
Two properties of the loop shape everything else. First, the API keeps no state between requests, so the conversation you resend IS the agent's working memory, and it grows every turn. Second, a production loop is bounded: it ends on Claude's final answer, with a turn or spending cap as a safety net, such as the Agent SDK's max_turns and max_budget_usd options. Every other pattern attaches to this loop at one specific point.
| Pattern | The problem it solves | Where it attaches to the loop |
|---|---|---|
| Tool-use loop | One call cannot act or look anything up | It is the loop |
| Subagents | The main conversation fills with raw material | A tool call that starts a separate loop with a fresh context |
| Memory | Each new session starts blank | A tool whose store lives outside the conversation |
| Context-window management | The resent conversation grows every turn | Rules for what each request carries and for how long |
Memorise the middle column: a question on this skill usually describes the problem and asks for the pattern that fits.
1.3.3 Subagents: send out readers, get back findings
The first version of the assistant did all the reading itself. Every document's full text stayed in the conversation for every later turn, so by the eightieth document the bill had climbed and answers missed facts plainly in the file. That last symptom has a name, context rot: as the number of tokens in the context grows, the model recalls information from it less accurately.
A subagent is the fix. It is a separate agent instance, started by the main agent (its parent), with its own fresh context: none of the parent's conversation, only its own instructions and the task you pass it. It does the reading (and any tool calls of its own) and returns only its final message. The raw text never reaches the parent's conversation. Think of a partner who sends two juniors into the archive: each comes back with a two-page memo, not with the boxes.
Here is a subagent written directly on the Messages API, so nothing is hidden. Look at the messages list, which carries no parent history, and at the return line, which hands back only the findings; load_text stands for your own document store.
def read_bundle(doc_ids: list[str], question: str) -> str:
"""Subagent: reads a few documents in a FRESH context, returns findings only."""
docs = "\n\n".join(f'<doc id="{d}">\n{load_text(d)}\n</doc>' for d in doc_ids)
reply = client.messages.create(
model=MODEL,
max_tokens=1500,
system="You review case documents for one question. Reply with at most "
"8 findings, each quoting the clause and citing its doc id.",
messages=[{"role": "user", # a NEW list: no parent history comes along
"content": f"{docs}\n\nQuestion: {question}"}],
)
return "".join(b.text for b in reply.content if b.type == "text") # findings only
# Main loop, when Claude requests the read_bundle tool:
findings = read_bundle(block.input["doc_ids"], block.input["question"])
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": findings}) # a page of findings, not the documents
This reader needs no tools, so one call is its whole loop; a subagent with tools runs a loop of its own. Three consequences follow. The subagent knows only what its prompt says, so the brief must carry the question, the document ids and the output format. Independent bundles can be read in parallel. And subagents are not free: each one makes requests of its own, so they pay off only when the material is bulky or the pieces are independent. The Claude Agent SDK packages the same pattern: subagents defined with a description, a prompt and, optionally, a restricted tool list.
1.3.4 Memory: what survives between sessions
On Thursday an associate opens the assistant again: "Carry on with the lease dispute." The new session knows nothing about Monday, because Monday's working memory was Monday's conversation. The tempting fix is to paste Monday's whole transcript in front of Thursday's first request. Resist it: you would pay again for every excerpt and tool result, and push the new session straight into context rot.
The pattern is memory outside the context window: the agent writes down what is worth keeping, and later reads back only what the current task needs. It is how the associate works: on Thursday nobody rereads 400 documents, only Monday's page of notes. One ready-made implementation is Anthropic's memory tool. You add {"type": "memory_20250818", "name": "memory"} to tools, and Claude can then view, create, edit, rename and delete files under a /memories directory. With the tool enabled, Claude checks that directory before it starts a task, so Thursday's session begins by reading Monday's notes.
The memory tool runs client-side, and that detail matters most. Claude only REQUESTS a memory operation; your application's handler, the code that carries out each request, executes it against storage you control and returns the result like any other tool result. The /memories path is only a prefix the handler maps onto real storage. Ingrid maps it to one folder per client matter, so notes from the lease dispute never surface in another client's session. That mapping is a confidentiality control, and it lives in her code, not in a prompt.
Two kinds of memory in one agent
Within a session
Across sessions
/memoriesBecause your code executes every memory operation, the safeguards are yours too. Validate every path, so /memories/../secrets.env cannot escape the directory; cap file sizes; expire notes nobody has opened in a long time. Claude usually declines to write sensitive information to memory, but "usually" is not a control: have the handler strip anything that must never be stored.
1.3.5 Context-window management: a budget, not a bucket
Subagents and memory help, but the loop's conversation still grows every turn. Everything in a request counts against the context window: the system prompt, the tool definitions, and every message, tool result and earlier reply. And quality does not wait for the hard limit: context rot sets in as the conversation grows. So context management belongs in the design from the first loop, not in a clean-up after things break.
Anthropic describes context as an attention budget that every token draws on, and the word is well chosen. A bucket works fine until it overflows; a budget shrinks with every purchase, so each token you add leaves less attention for the ones that matter. Anthropic's guiding principle is to find the smallest set of high-signal tokens that gets the job done. In practice you decide, for each kind of content, how long it stays in the conversation, using five techniques. Recognise each by the pressure it relieves.
| Pressure on the window | Technique | What it does |
|---|---|---|
| Old tool results nobody needs again | Tool-result clearing | Removes stale results from the history |
| A long conversation nearing the limit | Compaction | Replaces older turns with a summary and carries on |
| Bulky material needed for one step | Subagents | Read it in a separate context, return findings |
| Facts needed in a later session | Memory | Moves them out of the window into a store |
| Large sources that may never be needed | Just-in-time retrieval | Keeps ids and paths, loads content only when needed |
Anthropic's guidance matches three of them to kinds of task: compaction suits long back-and-forth conversations, note-taking memory suits work with clear milestones, and subagents suit research where parallel exploration pays off.
Here is the assistant's plan. Search returns ids and titles rather than full text, reading happens in subagents, confirmed dates go to the matter's memory folder, and a very long session gets compacted rather than cut off. Anthropic's context-engineering guidance names legal and finance work as likely candidates for a hybrid: load a little up front for speed (here, the matter summary) and let the agent retrieve the rest just in time.
1.3.6 Frameworks: what they package and what they hide
With the patterns clear, Tomasz raises the other half of this skill: build directly on the Claude API with the official client SDK (the anthropic package), or on a framework such as LangGraph? An agentic abstraction framework, or agent framework for short, is a library that packages the loop and the patterns around it, for agents and for fixed multi-step workflows alike. In Anthropic's summary, frameworks simplify low-level work such as calling the model, defining and parsing tools, and chaining calls together, which makes it easy to get started.
| Framework | Its central idea | What it hands you |
|---|---|---|
| LangGraph (from LangChain) | An agent or workflow as a graph of steps with shared state | Explicit branches, saved checkpoints, pauses for human review |
| Strands Agents (from AWS) | Model-driven: give it a model, a prompt and tools, and it runs the loop | A short path from tools to a working agent |
| PydanticAI (from the Pydantic team) | Typed Python agents | Validated structured outputs and tool arguments |
| Claude Agent SDK (from Anthropic) | Claude Code's agent loop as a library | Built-in tools, subagents, automatic compaction |
Recognise each by its central idea; underneath, each runs the same cycle: request, tool call, result, repeat.
The same summary carries a warning. Frameworks add layers of abstraction that can hide the actual prompts and responses, which makes an agent harder to debug, and they tempt you to add complexity a simpler setup did not need. Anthropic's advice is to start with the API directly, since many patterns take a few lines of code. If you adopt a framework, understand what it does underneath: wrong assumptions about what is under the hood are a common source of errors.
The team hit exactly that. Tomasz's LangGraph prototype worked in the demo, but each run cost several times more than Ingrid's hand-written loop. Logging the raw requests showed why: the graph's shared message list kept every document the reading step had loaded, so each later model call received all of it. The fix was a reading step with its own context that returns findings only. The framework was not wrong; the assumption about what it passed along was.
What a framework takes over, and what stays yours
The framework handles
You still own
A framework helps when its central idea matches a real need: branching workflows with checkpoints and human approval, a team already standardised on one, or several model providers behind one interface. The practical test is whether you can print the exact request it sends to Claude on any turn. If you can, the framework is a convenience; if you cannot, it is a black box. And as you move to production, Anthropic's advice is not to hesitate to reduce layers and build with basic components.
1.3.7 The exam traps
Most mistakes on this skill reach for brute force, or treat a pattern or a framework as magic.
- ✗ Pulling every document into the main conversation so that nothing gets lost. ✓ Keep pointers there instead: search returns ids and titles, subagents read the bulk and return condensed, cited findings, and stale results get cleared. Raw text in the main conversation is paid for on every later turn and invites context rot.
- ✗ Giving a subagent a one-line task and expecting it to know the case. ✓ Put the question, the inputs and the output format in its prompt. It starts fresh and sees nothing of the parent's conversation.
- ✗ Replaying every earlier transcript so the agent "remembers" a user or a matter. ✓ Keep curated facts in a memory store outside the window, scoped per user or matter, and read back what the step needs.
- ✗ Buying a bigger context window, or a bigger model, to cure drift. ✓ Control what enters the window. More tokens bring more rot, not more focus.
- ✗ Adding subagents or a framework because they sound advanced. ✓ Add each pattern for a problem you can name, starting from the plainest loop that works. Subagents add requests of their own; frameworks add layers.
- ✗ Treating the framework as a black box. ✓ Log the requests it actually sends. Wrong assumptions about what is under the hood are a common source of agent bugs.
Four tempting fixes for a drifting agent, one real one
1.3.8 Put it together: build a case-file assistant and break it
You now have every piece: the loop, subagents, memory and a managed context budget. A framework may package them or your own code may, but the mechanics are the same. To make them stick, build a small version, then remove one pattern at a time.
Context engineering (6.1) takes compaction and tool-result clearing into their full mechanics. Tool implementation (8.1) covers how to design and describe the tools this loop calls. And Claude hooks (7.3) are code that runs inside the loop to block a destructive action before it happens.
Key takeaways
- ✓ An agent is Claude using tools in a loop: Claude requests a tool, your code runs it and appends the result, until Claude answers or a safety cap ends the run.
- ✓ Subagents, memory and context-window management each attach to that loop, and you choose them by the problem you see.
- ✓ A subagent has a fresh context without the parent's conversation and returns only its final message, so brief it fully and use it to keep bulky work out of the main one.
- ✓ Memory across sessions lives outside the window; with the memory tool Claude requests each operation and your code executes it, so scoping and path checks are yours.
- ✓ The context window is a budget: everything counts, quality drops as it fills, and you decide what to clear, compact, isolate, store or load on demand.
- ✓ Frameworks such as LangGraph, Strands and PydanticAI package the loop and the patterns; start simple, adopt one when its central idea fits, and keep the actual requests visible.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
24 CCDV-F questions on Domain 1, free
Every question in the bank is tagged to a domain, so you can drill 24 questions on Agents and Workflows alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 1 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.