Home › Study guides › CCDV-F › Domain 2 › Lesson 2.5
CCDV-F · Domain 2 · 33.1% of the exam · Lesson 2.5 · 24 min read
Designing Claude applications: what the model actually sees
Why one prompt behaves differently in claude.ai, Claude Code and the API, and how to design content boundaries, schemas, sessions and plugins on purpose.
Written against skill 2.5 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
2.5.1 Why the same prompt behaves differently in your product
Tallyforge sells a web analytics dashboard to online shops, and its customers keep asking the same question about their weekly reports: why did this number move? So Soren, a back-end developer, is building a "report explainer". It is a panel beside each report where a customer's user types a question and gets a plain-language answer from Claude through the API. He tuned the prompt in claude.ai first, and there it was excellent. It knew which week "last week" meant, used the metric definitions he had uploaded to a project, and wrote tidy bullet points.
In the product, the same prompt misbehaved. It guessed at the dates, explained "sessions" with a textbook definition instead of Tallyforge's, and filled the panel with Markdown asterisks that the panel does not render. Worse, in a test with two demo customers, an answer for one mentioned a campaign that belonged to the other. Meanwhile Tallyforge's own staff use claude.ai and Claude Code every day with a shared team plugin, and they cannot see why the product should behave any differently.
Nothing is wrong with the model. It is the same Claude everywhere. What differs is everything wrapped around your words before they reach it: a system prompt (standing instructions placed before the conversation), tools, memory, files and earlier turns. A claude.ai chat arrives with a rich package already assembled; the API sends only what your code puts in the request. Claude application design is the discipline of deciding that package on purpose: which instructions, which data, in what shape, for which user, and for how long.
One prompt, two packages
The prompt in claude.ai
The same prompt through the API
2.5.2 What each interface wraps around the model
Here is the question that trips people up, and the exam likes it: if every interface runs the same model, why does a prompt that works in one fail in another? Because the model never receives "your prompt" on its own. It receives a request that an interface assembled, and each interface assembles a different one.
Picture a consultant who works at several client offices. At one, reception hands over a briefing pack, today's schedule and notes from the last meeting. At another, the consultant gets a single question on a sticky note. Same person, very different answers; only the pack changed. Precisely: every interface decides the system prompt, the tools, the memory and the conversation history that travel with your words.
| Interface | What it adds around your words | What that means when a prompt moves |
|---|---|---|
| claude.ai (web and mobile) | Anthropic's own system prompt (the current date, formatting habits), your account-wide instructions, project instructions and files, memory of past chats, tools such as web search | A prompt tuned here leans on all of this without you noticing |
| Claude Desktop | The same account features, plus desktop extensions: local Model Context Protocol (MCP) servers, small programs that give Claude tools for files and apps on your computer | Anything that relied on a local extension is missing elsewhere |
| Claude Code | Its own coding-agent system prompt, built-in tools to read and edit files and run commands, CLAUDE.md instruction files (delivered as a user message after that system prompt), auto memory (notes Claude Code keeps for itself), enabled plugins | All of it shapes Claude Code sessions; none of it reaches your product |
| Messages API and the client SDKs | Only what your request carries: your system field, your messages and your tools, plus the tool and format instructions the API builds from your own definitions |
Stateless (nothing carries over between requests): no system prompt, memory or date unless you add them |
| Claude Agent SDK | The engine behind Claude Code as a library, with a minimal system prompt unless you choose the claude_code preset or write your own; by default it also loads CLAUDE.md from the working directory and from ~/.claude/ |
Choose the system prompt, and whether CLAUDE.md loads, on purpose |
Memorise the API row: the request is the whole world. The client SDKs for Python, TypeScript and other languages add types, streaming helpers, retries and error handling, but nothing the model reads. Recognise the other rows well enough to spot what a prompt was quietly relying on.
That is exactly Soren's problem. In claude.ai, Anthropic's system prompt gave Claude the date, his project supplied the metric definitions, and the chat window rendered the Markdown. His API request carried none of it. The fix is not a cleverer sentence. It is a system prompt that writes down what the old prompt relied on: a role, today's date, Tallyforge's metric definitions, and an answer format the panel can display.
The staff side is the mirror image. Adaeze, a platform engineer, maintains the team's Claude Code plugin, and a CLAUDE.md file in the repository sets the coding conventions. Both shape her colleagues' Claude Code sessions; neither reaches the explainer's API calls. If the product needs a rule, the product's own request must carry it.
2.5.3 Content boundaries: what gets in, and what gets out
Soren's first request was one long string: his instructions, then the report data, then the user's question. It worked until a customer's report included a campaign named "IGNORE - internal test". The explanation left that campaign out of every total. Nobody attacked anything. The model had no way to tell where Tallyforge's instructions ended and the customer's data began, so it read a campaign name as an order.
A content boundary is the line your application draws around each kind of content: what enters the request, in which place, and what leaves it for the user. On the way in, three kinds of content matter.
- Instructions. Written by your team, reviewed like code, and placed in the system prompt. They say what the model is for and how to treat everything else.
- Data. The report, the user's question and any tool results. It changes on every request and often comes from people you do not control, so it travels in the conversation, never in the system prompt. It goes in the user turn inside labelled tags such as
<report>and<question>, or in a tool result. Anthropic's prompting guide recommends a tag per kind of content because it reduces misinterpretation. - Everything else. Other customers' reports, internal notes, secrets. In a product shared by many customers, each customer is a tenant, and one tenant's data has no place in another tenant's request. The task does not need any of this, so the request never carries it.
Three kinds of content, three decisions
Instructions
Data
Never sent
The boundary runs the other way too. What the model writes goes back through your code before a customer sees it, and your code decides what to display: the fields the panel expects, not raw output. The strongest output boundary is the input one. A model cannot reveal a report that was never in its request. No instruction such as "never mention other customers" is as reliable as leaving them out.
Text written deliberately to hijack a model, known as prompt injection, is a security problem with defences of its own. The design job here comes first: make sure the model can always tell your instructions from the material it works on.
2.5.4 Schema design: contracts for inputs, outputs and tools
The explainer panel shows an arrow badge (up, down or flat) next to each answer. Soren's prompt said "start your answer with the trend". The badge code then met "Up", "increasing", "Sessions rose modestly" and, once, a polite paragraph with no trend at all. Each was reasonable English and a bug in the product.
The fix is a schema: a machine-readable contract for the shape of data crossing a boundary. A Claude application has three of them, and each has its own design rule.
| Contract | Where it lives | Design rule |
|---|---|---|
| Input data | The user turn, inside labelled tags | Send structured data (JSON or a table) with units, the date range and the metric names your system prompt defines; the model cannot guess which week a figure covers or that "sessions" exclude bots |
| Model output | output_config.format with a JSON schema |
Required fields, an enum for every closed set, and additionalProperties: false so no surprise fields appear |
| Tool inputs | Each tool's input_schema, with strict: true where a malformed call would break something |
A detailed description of what the tool does and when to use it, and typed, required parameters |
Here is the explainer's output contract. Look at the enum on trend, a fixed list of allowed values that turns free text into a closed set, and at additionalProperties, which keeps unexpected fields out of the panel.
explanation_schema = {
"type": "object",
"properties": {
"summary": {"type": "string"},
"trend": {"type": "string", "enum": ["up", "down", "flat"]}, # a closed set, not prose
"drivers": {"type": "array", "items": {"type": "string"}},
"caveat": {"type": "string"}, # optional: gaps in the data
},
"required": ["summary", "trend", "drivers"],
"additionalProperties": False, # no surprise fields
}
response = client.messages.create(
model=MODEL, max_tokens=1024, system=SYSTEM_PROMPT, messages=history,
output_config={"format": {"type": "json_schema", "schema": explanation_schema}},
)
With structured outputs, the API constrains Claude's reply to this schema, so the badge code receives JSON with one of three trend values. Structured outputs also insists on additionalProperties: false for every object, so that line is a requirement as well as a good habit. The documentation names the exceptions: a refusal or a reply cut off at max_tokens may not match, and the capitalisation of an enum value is not guaranteed. Your code still checks what arrives.
Tool schemas work the same way. A get_report tool with a one-line description and an untyped id invites guesses; a detailed description and a typed, required report_id make the right call the easy one. Setting strict: true on a tool gives its inputs the same guarantee that output_config gives the reply.
2.5.5 Session hygiene: what a conversation carries, and for whom
Now the leak. In testing, a user at the demo shoe shop asked the explainer "Which campaigns have we discussed?", and the answer named a campaign from the demo bakery. The model did not remember the bakery. The Messages API is stateless: each request carries the full conversation, and the model knows only what is in it. If a bakery campaign appeared, Soren's code had put it there.
It had. For the prototype, he stored the conversation in one list at the top of the module. Every request from every customer appended to it, and every request sent all of it. It is a notepad on a shared reception desk: each visitor adds a line and can read every line above. A session is the history your application keeps and resends for one piece of work. Session hygiene is deciding what a session carries, whom it belongs to, and when it ends.
The fix fits in a few lines. Look at the key: each history belongs to one tenant, one user and one conversation, and each request sends only that history. Your server takes the tenant and user from the signed-in session, never from text the browser sends.
histories = {} # in production: a table keyed the same way, never one shared list
def ask(tenant_id, user_id, conversation_id, question, report_json):
key = (tenant_id, user_id, conversation_id) # SCOPE: one tenant, one user, one task
history = histories.setdefault(key, [])
turn = f"<question>{question}</question>"
if not history: # a new conversation starts with its own report
turn = f"<report>{report_json}</report>\n{turn}"
history.append({"role": "user", "content": turn})
response = client.messages.create(
model=MODEL, max_tokens=1024, system=SYSTEM_PROMPT,
messages=history, # this conversation's turns only
)
history.append({"role": "assistant", "content": response.content})
return response
Scoping answers "for whom". Hygiene also answers "for how long". A new report or an unrelated question starts a new conversation, because turns about another report now mislead more than they help. A long conversation about one report gets trimmed or summarised before each call rather than resent whole; doing that well is a context-management skill of its own. And a stored history is customer data, kept only as long as the product needs it.
The same rules apply to the staff's own tools. Each new Claude Code session begins with a fresh context window (everything the model can see at once), and CLAUDE.md files and auto memory are what carry knowledge into it. Within a session, context piles up. When Adaeze moves from one customer's data import to another's, the first customer's error logs and column names are still in context, steering the second job. The /clear command starts fresh between unrelated tasks, and /compact replaces a long history with a summary when she needs to keep going on the same one.
2.5.6 Plugin management: packages that run in every session
Tallyforge's staff side has its own design problem. Adaeze's team plugin is called tallyforge-dev. Over a few months colleagues also installed five more plugins "just in case". Sessions felt heavier, one plugin ran a hook nobody had read, and one Monday the team plugin behaved differently without anyone deciding it should.
A Claude Code plugin is a directory of components that Claude Code installs and loads as one unit; the table below names the main ones. A manifest at .claude-plugin/plugin.json usually names it, and its skills run with that name as a prefix, such as /tallyforge-dev:explain-metric. The same plugin format also installs on claude.ai, where a different set of components loads. None of it reaches a product's API calls.
The part that surprises people: an enabled plugin is part of every session, not only the sessions where you use it.
| Component | What it adds to every session where the plugin is enabled | What to review before enabling it |
|---|---|---|
| Skills (instructions Claude loads when relevant) and agents (helpers Claude can hand work to) | For those Claude can invoke on its own, the name and description sit in context every turn; the full text loads only when used | Whether the team needs them, since idle descriptions still cost context |
| Hooks | Shell commands that run at points such as after every edit, with your user permissions | The command each hook runs, in hooks/hooks.json |
| MCP servers | Tool servers that run alongside the session and give Claude their tools | Each server's command or URL, in .mcp.json |
| Manifest | The plugin's name, plus an optional version and description | Who publishes it, and which marketplace it comes from |
Plugin management comes down to three habits. Choose a plugin for a job the team actually does, and disable idle ones: claude plugin disable stops a plugin without uninstalling it. Review what the hooks and servers run before enabling anything, because what a plugin runs, it runs as you. Maintain it like code.
Adaeze installs the team plugin at project scope, which records it in the committed .claude/settings.json so it is enabled for everyone in the repository; each colleague still installs it once. Updates come from the team's marketplace, the catalog that lists the plugin and where to fetch it. Auto-update was on for that marketplace and the manifest set no version, so every commit counted as a new release. That was the Monday surprise: an unreviewed change reached everyone at their next session. Now changes go through review, and the manifest pins a version that Adaeze raises only for a release; until she does, everyone stays on the same copy.
2.5.7 The exam traps
Every trap in this skill comes from forgetting that the model sees only the package. Each wrong practice below looks reasonable on its own; each right one changes what the application assembles.
- ✗ Assuming a prompt, a CLAUDE.md or a team plugin that works in one interface carries over to your product. ✓ Rebuild what it relied on (role, date, definitions, format, tools) in the product's own system prompt and request. The API sends only what your code puts in it.
- ✗ Concatenating instructions, data and the user's question into one string. ✓ Put instructions in the system prompt and each kind of data in its own labelled tag, so the model can tell what to follow from what to read.
- ✗ Telling the model to keep other customers' data secret. ✓ Never send it. The request your code builds is the boundary; an instruction is not a control.
- ✗ Asking for "JSON, please" and parsing whatever comes back. ✓ Define a schema with required fields, enums for closed sets and
additionalProperties: false, and send it with structured outputs or strict tools. - ✗ Sharing one conversation history across users, or reusing it across unrelated tasks. ✓ Key every history by tenant, user and conversation, start fresh for new work, and
/clearbetween unrelated tasks in Claude Code. - ✗ Installing plugins "just in case" and letting them change without review. ✓ Enable what the team uses, read what its hooks and servers run, and release updates deliberately.
Four tempting fixes, one real one
2.5.8 Put it together: build a report explainer and break it
You now have every piece: interfaces, boundaries, schemas, sessions and plugins. The quickest way to make them stick is to build a small explainer, then break it twice and watch it fail the way Soren's prototype did.
The artefacts you just made are the start of the rest of the domain. Configuration management (2.6) puts prompts, CLAUDE.md, settings.json and plugin dependencies under version control, with pinned model versions, so a Monday surprise becomes a reviewed release. Context engineering (6.1) shows how to prune and compact the histories you now scope correctly. Output handling (6.3) validates what comes back even when a schema constrains it. AI application security (7.1) hardens the content boundary against text written to attack it.
Key takeaways
- ✓ The model responds to the package an interface assembles (system prompt, tools, memory and history), so the same prompt can behave differently in claude.ai, Claude Code and the API.
- ✓ The Messages API sends only what your request carries; rebuild whatever a prompt relied on elsewhere, and remember that CLAUDE.md and plugins shape Claude Code sessions, not your product.
- ✓ Keep trusted instructions in the system prompt and per-request data in labelled tags, and leave out anything the task does not need, because what is never sent cannot leak.
- ✓ Design a schema for every boundary: structured input data, a JSON output schema with required fields, enums and
additionalProperties: false, and tools with detailed descriptions and typed inputs. - ✓ Scope every conversation history to one tenant, one user and one task, start fresh for unrelated work, and use
/clearfor the same reason in Claude Code. - ✓ An enabled plugin is present in every session, so enable only what the team uses, review its hooks and servers, and release its updates deliberately.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
51 CCDV-F questions on Domain 2, free
Every question in the bank is tagged to a domain, so you can drill 51 questions on Applications and Integration alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 2 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.