Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 2 › Lesson 2.5

CCAR-P · Domain 2 · 13% of the exam · Lesson 2.5 · 23 min read

Prompt reuse: caching, modular prompts and Agent Skills

How to cache a repeated prompt prefix, keep shared rules in owned and versioned modules, and package know-how as Skills that load only when needed.

Written against objective 2.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.5.1 Why a copied prompt gets expensive, stale and stuck

Brightcairn Advisory, a global consulting firm, gives its consultants a proposal-writing assistant built on the Claude API. Twelve practices use it, each with its own prompt. Every one of those prompts opens the same way: a 9,000-token brand and compliance guide covering house style, the conflict-of-interest disclaimer and what no proposal may promise. Below it sits the practice's method, 15,000 to 25,000 tokens of phases, templates, fee tables and case studies. At the bottom comes the consultant's brief.

Delphine, the architect who has just inherited it, finds three problems that look like one. The first is cost and waiting time: about 6,000 requests a working day each resend the same 9,000 tokens, 54 million tokens a day processed from scratch before a draft can start. The second is drift. When Cormac, the head of risk and compliance, reworded the disclaimer in the spring, the change reached seven of the twelve prompts. The third is know-how stuck in place: each method lives inside one prompt, consultants who draft in the Claude app keep their own copies, and a renewal letter carries a case-study library it never uses.

A bigger model or a bigger context window fixes none of this. These are three kinds of waste, each with its own kind of reuse: prompt caching for the processing, modular prompts for the text and Agent Skills for the know-how.

Three things you can reuse

Prompt caching

Reuses processingof an identical prefix
Solves cost and latency
Needs a byte-identical prefix

Modular prompts

Reuses source textone owned module
Solves drift and upkeep
Needs an assembly step

Agent Skills

Reuses know-howloaded on demand
Solves know-how stuck in one prompt
Needs a code execution environment
Each mechanism reuses something different, so each solves a different problem; Brightcairn's assistant needs all three.

2.5.2 Prompt caching: pay once for the prefix you repeat

Here is the belief that trips people up: surely the API remembers a guide it receives thousands of times a day? It does not, unless you ask. Each request is processed from its first token, so the guide is read, and billed, afresh every time. With prompt caching, you mark a point in the request with cache_control, a cache breakpoint, and the API keeps the processed state of everything up to it for a few minutes. A later request that begins with exactly the same content reads it back at a fraction of the input price, and its first token arrives sooner.

The catch is in "exactly the same". The cache key covers the whole prefix in a fixed order, tools, then system, then messages, up to and including the block with the breakpoint. Picture an interpreter who has studied your 60-page briefing pack and marked where it ends. Hand them the same pack and they go straight to today's question; change one word on page 2 and the mark is worthless. Precisely: a request reads an entry only when everything up to and including the marked block is byte-identical, so one changed character invalidates every breakpoint after it.

So the layout rule is static content first, varying content last. Three practices opened their system prompt with "Proposal for {client}, prepared {date}", so no two requests began alike and nothing could ever be read from cache. Delphine moved the client, the date and the brief into the user message. Then she set two breakpoints: one after the shared guide, identical for all twelve practices, and one after each practice's method. A request can carry up to four; they cost nothing and let parts that change at different speeds be cached separately.

Where Brightcairn's breakpoints go

TOOLSsame list, same order, every request
SHARED GUIDE9,000 tokens, all practices; breakpoint 1
PRACTICE METHODsame within a practice; breakpoint 2
THE BRIEFclient, date, request: never cached
Everything before a breakpoint must be identical to be read from cache, so what changes on every request goes after the last one.

In code, look at the two cache_control markers, each on the last block of a stable part, and at the user message, which carries everything that varies.

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=8000,
    tools=PROPOSAL_TOOLS,                          # same list, same order, every request
    system=[
        {"type": "text", "text": SHARED_GUIDE,     # 9,000 tokens, identical for all practices
         "cache_control": {"type": "ephemeral"}},  # breakpoint 1
        {"type": "text", "text": practice_method,  # identical within one practice
         "cache_control": {"type": "ephemeral"}},  # breakpoint 2
    ],
    messages=[{"role": "user",                     # everything that varies goes here
               "content": f"Client: {client_name}\nDate: {today}\n\n{brief}"}],
)
u = response.usage
log(u.cache_creation_input_tokens,                 # written to the cache by this request
    u.cache_read_input_tokens,                     # read from the cache: the saving
    u.input_tokens)                                # after the last breakpoint, full price

Automatic caching, a single cache_control at the top level of the request, instead puts the breakpoint on the last cacheable block. That suits a growing chat. Here, though, the last block is the brief, so every request would write a new entry and none would read one.

2.5.3 Lifetime, price and what breaks a cache

The layout decides whether a hit is possible; three facts decide whether caching pays. The first is lifetime. An entry lives five minutes and every hit refreshes it free, so Brightcairn's guide, read every few seconds, stays warm all day. A one-hour lifetime ("ttl": "1h") wins when a prefix returns less often than every five minutes but within the hour, such as a partner who resumes a 40,000-token tender after a 25-minute call. The clock starts when a request begins, not when its reply ends.

The second is price, in multiples of the normal input price. A write costs 1.25 times for the five-minute cache and 2 times for the one-hour cache. A read costs 0.1 times on most models, and less on the newest top tiers: 0.05 times on Claude Opus 5.5 and 0.025 times on Claude Fable 5.1. So a five-minute entry pays for itself on its first hit, and a one-hour entry on its second. On most models, reads don't count toward your input-token rate limit either. Caching never changes the reply.

The third is the minimum length. A shorter prefix is processed without caching, and without any error. The minimum is 512 tokens on Claude Sonnet 5.5 and Claude Opus 5.5, 1,024 on Claude Sonnet 5 and 4,096 on Claude Haiku 4.5. The only evidence is in usage: if cache_creation_input_tokens and cache_read_input_tokens are both zero, nothing was cached. The table shows what else breaks a cache, layer by layer.

What changes What it invalidates What the architect does
Tool definitions: added, removed, reordered or reworded Everything: tools, system and messages Send one tool list in a fixed order, serialised the same way every time
The system prompt, by even one character System and messages; tools stay cached Keep dates, IDs and names out of it; release module changes deliberately
The model Everything, because each model has its own cache Keep traffic that shares a prefix on one model
The list of Skills attached to the request The cached prefix, since Skills render into the system prompt Attach one list, in one order, at pinned versions
tool_choice, or adding or removing images Messages Keep these settings the same across requests
Thinking settings or effort level Messages, and on some models system and tools too Set them per workload, not per request

Memorise the layer order and the first two rows; recognise the rest. When an expected hit goes missing, cache diagnostics names the cause: send the previous response's id in a diagnostics object, and the API reports where the two requests first differ, such as system_changed. On the Claude API, caches are also isolated per workspace, a subdivision of an organisation's account, so twelve practice workspaces would hold twelve caches of one guide.

2.5.4 Modular prompts: one source of truth for shared rules

Caching made requests cheaper. It did nothing for the disclaimer that reached only seven prompts, because that is a maintenance failure. Twelve copies mean twelve places to edit and no way to prove which wording a proposal was drafted under. The fix is the one you would apply to duplicated code: extract the shared parts into prompt modules, give each an owner and a version, and assemble every prompt from them.

Delphine split the guide by owner and rate of change. Brand voice belongs to marketing and changes yearly; output formats belong to the proposals office; compliance rules belong to Cormac's team and change most months. Each module lives in version control, changes through its owner's review, and ships as a numbered version once the regression evals pass. A manifest per use case lists which modules go in, at which versions and in which order. Look at the pinned versions and where the breakpoints fall.

use case: proposal-writer / practice: public-sector
1. brand-voice        v4   owner: Marketing          tag: <brand_voice>
2. output-formats     v7   owner: Proposals office   tag: <output_formats>
3. compliance-rules   v19  owner: Risk & compliance  tag: <compliance_rules>
-- cache breakpoint 1: identical for all 12 practices up to here --
4. practice-method    v11  owner: Public-sector lead tag: <practice_method>
-- cache breakpoint 2 --
request (not a module): client, date and brief, in the user message
release rule: a new module version ships only after the regression evals pass for all 12 practices

One module library, twelve assembled prompts

Module librarybrand voice, output formats, compliance rules
Public sectorbrand v4, formats v7, compliance v19
Healthcare strategysame shared modules, own method
Financial servicessame shared modules, own method
Nine more practicesone manifest each
Each shared module has one owner and one current version, and every practice's prompt is assembled from the library instead of carrying its own copy.

An assembler builds each prompt from its manifest, and it must be deterministic to stay cacheable: the same manifest yields the same bytes, with a fixed order, no generated timestamps and stable serialisation. Each module sits in its own XML tag, named the same way in every prompt, so Claude can tell brand rules from compliance rules. Delphine also logs the manifest version with every request, so any proposal's disclaimer is one lookup away, and a compliance change is one edit that reaches all twelve prompts at the next release.

2.5.5 Agent Skills: know-how packaged once, loaded when needed

Caching made the practice methods cheap to resend and modules made them maintainable, but each is still welded into one prompt. The public-sector method runs to 22,000 tokens, sixty case studies included, and its consultants keep drifting copies in the Claude app. A renewal letter needs none of the case studies and a new bid needs three, yet every request carries all sixty.

Agent Skills package that kind of know-how once. A Skill is a folder: SKILL.md at the top, with YAML front matter holding a name and a description, then instructions in Markdown, plus the references, templates and scripts those instructions point to. Anthropic's docs compare it to an onboarding guide for a new team member. What makes it more than a folder of text is progressive disclosure: Claude loads it in stages.

How a Skill loads

METADATAname and description: always, about 100 tokens
INSTRUCTIONSthe body of SKILL.md: when a task matches
FILESreferences and scripts: only when needed
Only the metadata costs context on every request; the instructions load when a task matches the description, and each file loads only when the instructions send Claude to it.

Claude reads the body of SKILL.md only when a request matches the description; the docs put that body at under 5,000 tokens. A bundled script runs without its code entering the context: only its output does. So the description does the routing, and it must say what the Skill does and when to use it. Look for both at the top of Brightcairn's public-sector Skill.

---
name: public-sector-proposals
description: Method, templates and fee rules for Brightcairn public-sector bids and renewals. Use when drafting or reviewing a proposal for a government body, agency or public institution.
---
# Public-sector proposals

1. Identify the procurement route from the brief: open tender, framework call-off or renewal.
2. Follow the phase plan for that route in `method/phases.md`.
3. Price with `scripts/fee_calculator.py`; never quote a day rate by hand.
4. For a new bid, choose at most three case studies from `case-studies/index.md`.
5. For a renewal, skip case studies and start from `templates/renewal.md`.

The same folder format works in the Claude apps, in Claude Code and on the API. On the API, an uploaded Skill is shared across its workspace. Each request names up to 20 Skills in its container parameter and includes the code execution tool: Claude reads and runs the Skill's files in its sandbox, which has no network access. In the apps, users can upload their own Skills once code execution is on, and Team and Enterprise owners can provision a Skill for the whole organisation. Claude Code reads folders: ~/.claude/skills/ for one person, .claude/skills/ in a repository, or a plugin. An API upload does not appear in the apps or in Claude Code, so the source lives in version control and each surface is deployed from it.

Delphine turned each method into a Skill, provisioned the twelve to consultants in the Claude app, and attaches all twelve to every assistant request in one order, at pinned versions. Changing the list, its order or, under latest, a Skill's description breaks the cache, so one fixed list keeps one shared prefix and lets a joint bid load two methods. Because twelve descriptions compete for Claude's attention, her evals check that each Skill triggers for its own practice and no other. The compliance rules stay out: a Skill loads only when Claude judges a task relevant, and those rules must govern every draft.

Treat a Skill like software you install, since it can carry scripts and instructions that steer tools: review every file, evaluate it before release, and keep the previous version for rollback. Check data handling too: prompt caching is eligible for zero data retention (ZDR), but Agent Skills are not covered by ZDR arrangements.

2.5.6 Choosing and combining the three

The requirements usually say which problem you face; the table turns that into a choice.

Option When it wins What it costs
Prompt caching A long prefix, above the model's minimum, repeats exactly within the cache lifetime, and the requirement is cost, latency or throughput Writes at 1.25 or 2 times the input price; a byte-identical prefix; one cache per workspace and model; no help with upkeep or reuse elsewhere
Modular prompts Several prompts share components that different teams own and change, and the requirement is consistency, auditability or one place to change An assembly step, a manifest and a release process, with evals on every assembled prompt; no token saving on its own
Agent Skills Procedural know-how, templates or scripts that several agents or surfaces reuse and only some requests need A code execution environment, metadata on every request, trigger accuracy to evaluate, a security review, no ZDR coverage

Read the third column as closely as the second: a workload under ZDR, a prefix below the minimum or a component nobody will own can rule out an option that wins on paper.

Brightcairn's final design combines all three. What makes it architecture rather than tuning is a written justification that partners and compliance can check. Here is the core of Delphine's decision record.

Decision record 14: prompt reuse for the proposal assistant. Owner: Delphine. Reviewers: Cormac, the twelve practice leads.
Context: 12 practices; a 9,000-token shared guide on about 6,000 requests a working day; old disclaimer wording live in 5 of 12 prompts; methods of 15,000 to 25,000 tokens, copied into the Claude app and mostly unused by any one request.
Decision 1: split the shared guide into brand-voice, output-formats and compliance-rules modules, each owned by one team and assembled from a per-practice manifest; log the manifest version on every request.
Decision 2: stable prefix first with a breakpoint at its end and the default five-minute lifetime, since the prefix is read every few seconds; client, date and brief after the last breakpoint.
Decision 3: practice methods become Skills built from one source: all twelve attached to the assistant in one fixed order at pinned versions, and provisioned to consultants in the Claude app; compliance rules stay in the system prompt because they must govern every draft.
Measures: share of input tokens read from cache; stale module versions in production (target zero); Skill trigger accuracy per practice from the eval suite.

2.5.7 The exam traps

Every trap here reaches for a mechanism that does not match the problem, or breaks the one condition a mechanism depends on.

  • ✗ Putting the date, a request ID or the user's name at the top of a prompt you want cached. ✓ Keep the prefix byte-identical and put everything that varies after the last breakpoint; one changed character invalidates all that follows it.
  • ✗ Cutting or summarising required policy text to save tokens. ✓ Keep it whole, put it first and cache it. Dropping required context saves money by making the answers worse.
  • ✗ Lengthening the cache lifetime to fix a near-zero hit rate. ✓ Find what differs between requests (tools, system prompt, model, Skills list) with the usage fields or cache diagnostics, and move the breakpoint to the last block that stays identical. No lifetime makes two different prefixes match.
  • ✗ Expecting caching to fix drift, or modules to cut the bill. ✓ Caching reuses processing and modules reuse source text. Name the problem, then pick the mechanism.
  • ✗ Packaging rules that must always apply as a Skill. ✓ A Skill loads only when a task matches it. Always-on rules belong in the system prompt.
  • ✗ Running production on latest, for a Skill or a prompt module. ✓ Pin versions, promote a new one only after the evals pass, and keep the previous one for rollback.

2.5.8 Put it together: cache, modularise and package one assistant

You now have every piece. To make it stick, watch a cache hit, break it with one timestamp, and watch a Skill lose its trigger.

Progressive discovery versus monolithic context (3.8) takes the Skills idea into integration design. Optimising token usage, latency and cost (4.5) turns the usage fields you logged into cost-per-request figures. Configuring Claude tools for teams (7.1) is where Skills reach developers through Claude Code projects and plugins.

Key takeaways

  • ✓ Reuse works at three levels: caching reuses processing, modules reuse source text and Skills reuse know-how, and the problem you name decides which one applies.
  • ✓ Caching reads back an identical prefix, in the order tools, system, messages, up to a cache_control breakpoint, so static content goes first and everything that varies after the last breakpoint.
  • ✓ Entries live five minutes, refreshed on each hit, or one hour; writes cost 1.25 or 2 times the input price, reads a tenth or less, and a prefix below the model's minimum is silently not cached.
  • ✓ A change to tools, system prompt, model or Skills list invalidates the cache from that layer on, so check the usage fields instead of assuming a hit.
  • ✓ Modular prompts give each shared component one owner and one version, assembled deterministically from a manifest and released only after evals pass.
  • ✓ A Skill is a SKILL.md folder reused across agents and surfaces and loaded by progressive disclosure; keep always-on rules in the system prompt, and pin Skill versions in production.
  • ✓ Justify each choice by the requirement it meets and the cost it brings, and combine the three when the problems coexist.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

24 CCAR-P questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 24 questions on Claude Models, Prompting & Context Engineering alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources