Home › Study guides › CCAR-F › Domain 5 › Lesson 5.1
CCAR-F · Domain 5 · 15% of the exam · Lesson 5.1 · 22 min read
Conversation context: keeping the facts that matter
Why summarising a long support chat blurs a $84.20 refund, and how a case-facts block, trimmed tool results and key findings placed first keep facts exact.
Written against task statement 5.1 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
5.1.1 How "$84.20 by Friday" became "a refund was discussed"
A customer opens a chat because order 5530, a kettle, arrived with a cracked base. At the twelfth turn the support agent confirms the fix: a refund of $84.20 to the original card, back by Friday. Then the conversation wanders, as real ones do: a second order running late, loyalty points, a new delivery address. Forty turns in, the customer returns to what they care about: "So the $84.20 will be on my card by Friday?" The agent replies, "I can see a refund was discussed. Could you remind me of the amount?"
Nothing crashed and no tool failed. The conversation grew long, so the application condensed the older turns into a summary to save space, and later summarised again, folding the old summary into a new one. Each pass kept the gist and shed a detail: "$84.20 refund for order 5530 by Friday" became "refund offered for the damaged kettle", then "a refund was discussed". The model answered faithfully from what it was given. The amount was no longer in it.
A model sees nothing of a conversation except the request your code sends. Everything in that request, the instructions, the conversation so far and every tool result, sits in the context window: all the text the model can see while it writes one reply. The window is measured in tokens, the small word pieces a model reads and writes, and it is large but finite. So in a long interaction something has to decide what goes in, in what form and in what order. That is context management, and this lesson is about doing it so the facts a customer depends on survive a long session.
5.1.2 The request is the only memory there is
Here is the belief behind most of the trouble in this lesson: that the model remembers the conversation. It doesn't. The Messages API (the application programming interface your code calls to reach Claude) is stateless: it keeps nothing between requests, so, as its documentation puts it, you always send the full conversational history. If you build on the Claude Agent SDK (a software development kit, a ready-made library for building agents), it keeps the history and resends it for you. The rule underneath is the same.
What the complete history buys is coherence. Suppose your code saves tokens by sending only the customer's newest message. The agent can no longer see that it verified the customer, that it apologised, or which order "it" means. So it asks for the order number again, or contradicts a promise it made ten minutes earlier. That looks like a model mistake. It is a missing-history mistake.
The price is growth. Each turn adds messages, and everything in the request, tool results included, counts toward the window. The docs also warn that as the token count grows, accuracy and recall degrade, a decline they call context rot. So a long session eventually has to be condensed, by your code or, in the Agent SDK, automatically as the window nears its limit. Condensing is where facts get lost; the rest of this lesson is about condensing without losing them.
| What each request carries | What the agent can do | What goes wrong |
|---|---|---|
| Only the newest message | Answer that one message | Asks for order 5530 again; forgets what it promised |
| The complete history, unmanaged | Stay coherent for a while | Grows every turn until it is summarised and the details blur |
| The history (condensed when long), a case-facts block, trimmed tool results | Stay coherent and exact | Nothing to fix: this is the design the lesson builds |
The first row is the shortcut to watch for: it looks like a token saving, and it costs the agent its coherence.
5.1.3 Summaries keep the gist and lose the numbers
Once the history is too long to send whole, the obvious move is the one from the opening story: replace the older turns with a summary, and later summarise the summary. That is progressive summarisation. It works, which is what makes it dangerous. A summary exists to keep the gist ("damaged item, refund offered"), and particulars are exactly what it is built to shed. Each pass paraphrases the last paraphrase, so the losses compound: no single summary looks wrong, yet the third no longer contains the $84.20.
Four kinds of detail blur first: numerical values, percentages, dates and customer-stated expectations. "$84.20" becomes "a refund", "we'll waive the 15% restocking fee" becomes "fees were discussed", and "by Friday" becomes "soon". The customer's own asks, such as "on my card, not as store credit", tend to vanish altogether, because summaries lean towards what the agent did rather than what the customer wanted.
The fix is structural. Pull the transactional facts out of the prose into a case-facts block: a short, structured record of amounts, dates, order numbers and statuses. Your code keeps it OUTSIDE the conversation history and includes it, whole, in every request, so the summariser never touches it. Your code also keeps it current. Facts from tools go straight in: when process_refund returns, "refund promised" becomes "refund issued". Facts stated in the chat, such as the request for a refund to the card, need an extraction step, for example a small Claude call that returns them in a fixed structure. The history can then be condensed as hard as you like, because nothing the case depends on lives only there.
Restaurants solved this long ago. A waiter relaying an order from memory turns "table six, steak medium-rare, no garlic, it's a birthday" into "table six, a steak". So the order goes on a ticket pinned to the rail, and every cook reads the ticket, not the retelling. Precisely: the case-facts block is copied into each request exactly as stored, so its facts are as exact on turn 60 as on turn 12.
Summary after three rounds versus summary plus a case-facts block
Summary alone
Summary plus case facts
| What goes in the block | Example from this chat | Why the summary cannot hold it |
|---|---|---|
| Order number | 5530 | "The order" is useless when the next tool call needs an ID |
| Amount | $84.20 refund | Paraphrase rounds or drops it: "about $80", "a refund" |
| Percentage | 15% restocking fee, waived | "Fees discussed" loses the rate and the decision |
| Date | Refund due by Fri 2 Oct | "Soon" cannot be kept or checked; "Friday" is ambiguous next week |
| Status | Refund promised, not yet issued | Promised versus issued decides the agent's next action |
| Customer-stated expectation | Back on the original card, not store credit | Summaries record the agent's actions, not the customer's asks |
Memorise the transactional core (amounts, dates, order numbers, statuses). Percentages and customer-stated expectations belong in the block too, because summaries blur them first.
Now the customer raises the late order. One flat block ("amount: $84.20, status: pending, due: Friday") cannot hold two issues without mixing them: which amount, which status? For multi-issue sessions, keep a separate context layer of structured issue data, with one entry per issue keyed by order ID. Your code adds an entry when a new issue comes up, and updates only that entry when its facts change. Here is the layer for this chat as your code might render it at the top of each request. Look at the two Order lines: every amount, status and date sits under the order it belongs to.
CASE FACTS (kept by the application, never summarised)
Order 5530 - kettle, cracked base
refund: $84.20 to the original card (customer: not store credit)
restocking fee: 15%, waived
status: refund promised, not yet issued
due: Fri 2 Oct
Order 5561 - late delivery
status: in transit, redelivery booked
due: Wed 30 Sep
5.1.4 Trim tool results before they pile up
The summary is not the only thing growing. Tool results are often the heaviest part of a support conversation. A lookup_order call against a real order system can return forty or more fields: warehouse codes, carrier IDs, tax breakdowns, internal flags, audit timestamps. To handle a return, the agent needs perhaps five: the order number, the items, the amount paid, the delivery date and whether the return window is still open.
Here is the part people underestimate. Because the history is resent on every request, a result, once appended, travels in every later request too. Three lookups put 120 fields into the history, about 105 of them noise, and they ride along until the session ends. They use up tokens on every turn, and they cost attention: the fields that matter sit among dozens that don't, and the conversation itself becomes a small part of what the model reads.
Think of a parcel that arrives in a box of packing foam. You shelve the item, not the foam; shelve the foam every time and within a week the shelf is full of it. In precise terms: trim verbose tool output to the fields the current task needs BEFORE the result is appended, because from then on every request carries whatever you kept.
Where should the cut happen? The support agent reaches your backend through custom tools served over the Model Context Protocol (MCP), the standard way to plug outside systems into an agent. You build those tools, so the cleanest place to trim is the tool itself: lookup_order returns a lean record shaped for returns.
If a tool is shared or not yours to change, trim in between. In the Agent SDK, a PostToolUse hook (your own code that runs after a tool returns and before Claude sees the result) can return updatedToolOutput, and Claude reads that trimmed version instead of the original. Either way the decision is made in code, for every result, not left to the model. If a later step turns out to need another field, add it to the list you keep.
Trimming a lookup before it enters the history
lookup_order returns40+ fields5.1.5 The middle of a long input gets the least attention
Even with every fact exact and every result trimmed, a long request has one more weakness, and it is about position. Models reliably use what sits at the beginning and the end of a long input, but may omit findings buried in the middle. The exam guide calls this the lost in the middle effect. The refund promise was made at turn 12 of 40. If the only record of it is in the history, it sits in exactly the stretch that gets the least attention, summarised or not.
You know this from a long list read aloud: the first items stick, the last items stick, and the middle blurs. Stated precisely: what a model uses from a long input depends partly on where each piece sits. So anything that must not be missed belongs at the start, clearly labelled, not wherever it happened to arrive.
The fix is two moves. First, put a key findings summary at the beginning of any aggregated input, meaning any input your code assembles from several parts. For the support agent, that is the case-facts block at the top of the request, above the summary, the lookups, the policy text and the recent turns. Second, organise the detail under explicit section headers ("Order 5530", "Order 5561", "Returns policy", "Conversation so far"), so the model can find each part and tell them apart.
Anthropic's long-context prompting tips point the same way, though they don't use the exam guide's name for the effect: long documents near the top, each labelled with its source, and your question at the end. In Anthropic's tests, a question at the end can improve response quality by up to 30 percent, especially with complex inputs made of several documents.
The same input, two layouts
Arrival order
Key facts first
The effect bites hardest where many results are combined, as in a research pipeline that hands ten subagent reports to one synthesis step. The reports in the middle are the likeliest to go missing, and the same fix applies: key findings first, one headed section per report.
5.1.6 Upstream agents: send facts, not transcripts
So far one agent has been managing its own context. When agents feed each other, the same problems arrive all at once, and a multi-agent research system shows them best. There, a coordinator delegates to four subagents: one searches the web, one analyses documents, one synthesises the findings and one writes the report. What the search and analysis subagents send back becomes the synthesis subagent's input. They are upstream, it is downstream, and its context budget, the room it has for input, is limited. Send it pages of narrative and reasoning, and the budget is spent before synthesis starts. Send it claims with no date or source, and nothing downstream can put them back.
Start with what a finding needs to be usable later. Suppose the search subagent reports that bus ridership rose 18% after a city made its buses free. The synthesis subagent needs to know WHEN, because a 2019 figure and a 2025 figure are different claims. It needs to know where (which document, which section) to cite it. And it needs to know how the figure was measured: a passenger survey and automatic passenger counters are not equally strong evidence. So you require subagents to include metadata in their structured outputs: dates, source locations and methodological context, captured when the finding is made, because nobody downstream can reconstruct them.
The second technique is about volume. When the downstream agent's context budget is tight, change the upstream agents to return structured data (key facts, citations, relevance scores) instead of verbose content and reasoning chains. The subagent still reasons at length, but in its own context; only the distilled result travels. Anthropic's context engineering article describes this shape: a subagent may explore using tens of thousands of tokens, yet return a condensed summary of often 1,000 to 2,000 tokens. It is the trimming rule applied to agents.
Here is one finding in the shape the synthesis subagent should receive. The lines that matter are source, published and method, the metadata, and relevance, which lets synthesis decide what to keep when its budget is tight.
{
"findings": [
{
"claim": "Bus ridership rose 18% in the first year after fares were removed",
"source": "https://example.org/transit-review.pdf, section 4.2",
"published": "2025-03-14",
"method": "Automatic passenger counters, one city, 12 months before and after",
"relevance": 0.92
}
]
}
5.1.7 The exam traps
Each trap below either leaves a fact to prose, to position or to a history cut short, or adds capacity where structure was needed.
- ✗ Letting the running summary carry amounts, dates and promises. ✓ Keep them in a case-facts block outside the summarised history, included in every request. A longer or more frequent summary only slows the blur.
- ✗ Sending only the latest message to save tokens. ✓ Send the complete conversation history with every request. The API remembers nothing, so a turn you leave out is a turn the agent never had.
- ✗ One flat list of facts for a customer with several issues. ✓ One structured entry per issue, keyed by order ID, so amounts and statuses cannot drift between orders.
- ✗ Appending the full 40-field
lookup_orderresult, or buying a bigger context window. ✓ Trim to the task's fields before the result is appended. A bigger window postpones the problem and still carries the noise on every turn. - ✗ Concatenating long inputs in arrival order. ✓ Key findings summary first, explicit section headers, the question last. Telling the model to "read carefully" does not change where things sit.
- ✗ Letting upstream agents return their full reasoning, or stripping dates and sources to save space. ✓ Structured key facts with citations, dates, method and relevance scores: what synthesis needs, without the reasoning it can do without.
5.1.8 Put it together: build a support context that keeps its facts
You now have every piece: complete history for coherence, and a facts block with one entry per issue for exact numbers. Add trimmed tool results, key findings first under headers, and structured upstream outputs that carry their metadata. The build below puts the support-side pieces into one small agent, then takes the facts block away so you can watch the numbers go.
Keeping facts exact is the footing for the rest of Domain 5. Escalation (5.2) decides when a case goes to a human, and a handover is only as good as the facts that travel with it. Error propagation (5.3) applies the structured-output idea to failures: what was tried and why it failed, not a bare "unavailable". Large codebase exploration (5.4) is the same context pressure inside Claude Code, handled with scratchpad files, subagents and /compact. Provenance (5.6) extends the metadata requirement into claim-source mappings that survive synthesis.
Key takeaways
- ✓ The API is stateless, so every request must carry the complete conversation history the agent needs; send less and the agent loses coherence.
- ✓ Progressive summarisation keeps the gist and blurs amounts, percentages, dates and customer-stated expectations, and each pass compounds the loss.
- ✓ Keep transactional facts in a persistent case-facts block included in every request outside the summarised history, with one structured entry per issue when a session has several.
- ✓ Tool results are resent on every turn after they are appended, so trim them to the fields the task needs before they accumulate.
- ✓ Long inputs lose their middle: put a key findings summary first, organise the detail under explicit section headers, and ask the question last.
- ✓ Subagents return structured findings with dates, source locations and method; when downstream budgets are tight, upstream agents return key facts, citations and relevance scores instead of reasoning chains.
Check your understanding
4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
54 CCAR-F questions on Domain 5, free
Every question in the bank is tagged to a domain, so you can drill 54 questions on Context Management & Reliability alone, or sit the full 60-question timed simulator.
Open the CCAR-F question bank → Back to Domain 5 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.