Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 2 › Lesson 2.5

CCDV-F · Domain 2 · 33.1% of the exam · Lesson 2.5 · 24 min read

Designing Claude applications: what the model actually sees

Why one prompt behaves differently in claude.ai, Claude Code and the API, and how to design content boundaries, schemas, sessions and plugins on purpose.

Written against skill 2.5 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.5.1 Why the same prompt behaves differently in your product

Tallyforge sells a web analytics dashboard to online shops, and its customers keep asking the same question about their weekly reports: why did this number move? So Soren, a back-end developer, is building a "report explainer". It is a panel beside each report where a customer's user types a question and gets a plain-language answer from Claude through the API. He tuned the prompt in claude.ai first, and there it was excellent. It knew which week "last week" meant, used the metric definitions he had uploaded to a project, and wrote tidy bullet points.

In the product, the same prompt misbehaved. It guessed at the dates, explained "sessions" with a textbook definition instead of Tallyforge's, and filled the panel with Markdown asterisks that the panel does not render. Worse, in a test with two demo customers, an answer for one mentioned a campaign that belonged to the other. Meanwhile Tallyforge's own staff use claude.ai and Claude Code every day with a shared team plugin, and they cannot see why the product should behave any differently.

Nothing is wrong with the model. It is the same Claude everywhere. What differs is everything wrapped around your words before they reach it: a system prompt (standing instructions placed before the conversation), tools, memory, files and earlier turns. A claude.ai chat arrives with a rich package already assembled; the API sends only what your code puts in the request. Claude application design is the discipline of deciding that package on purpose: which instructions, which data, in what shape, for which user, and for how long.

One prompt, two packages

The prompt in claude.ai

Anthropic's system prompttoday's date, formatting habits
Your instructions and project files
Memory of past chats
Your prompt

The same prompt through the API

A system promptonly if you write one
Reference dataonly what your code adds
Earlier turnsonly the history you resend
Your prompt
The words are identical, but the package around them is not. Through the API, the model sees only what your request carries.

2.5.2 What each interface wraps around the model

Here is the question that trips people up, and the exam likes it: if every interface runs the same model, why does a prompt that works in one fail in another? Because the model never receives "your prompt" on its own. It receives a request that an interface assembled, and each interface assembles a different one.

Picture a consultant who works at several client offices. At one, reception hands over a briefing pack, today's schedule and notes from the last meeting. At another, the consultant gets a single question on a sticky note. Same person, very different answers; only the pack changed. Precisely: every interface decides the system prompt, the tools, the memory and the conversation history that travel with your words.

Interface What it adds around your words What that means when a prompt moves
claude.ai (web and mobile) Anthropic's own system prompt (the current date, formatting habits), your account-wide instructions, project instructions and files, memory of past chats, tools such as web search A prompt tuned here leans on all of this without you noticing
Claude Desktop The same account features, plus desktop extensions: local Model Context Protocol (MCP) servers, small programs that give Claude tools for files and apps on your computer Anything that relied on a local extension is missing elsewhere
Claude Code Its own coding-agent system prompt, built-in tools to read and edit files and run commands, CLAUDE.md instruction files (delivered as a user message after that system prompt), auto memory (notes Claude Code keeps for itself), enabled plugins All of it shapes Claude Code sessions; none of it reaches your product
Messages API and the client SDKs Only what your request carries: your system field, your messages and your tools, plus the tool and format instructions the API builds from your own definitions Stateless (nothing carries over between requests): no system prompt, memory or date unless you add them
Claude Agent SDK The engine behind Claude Code as a library, with a minimal system prompt unless you choose the claude_code preset or write your own; by default it also loads CLAUDE.md from the working directory and from ~/.claude/ Choose the system prompt, and whether CLAUDE.md loads, on purpose

Memorise the API row: the request is the whole world. The client SDKs for Python, TypeScript and other languages add types, streaming helpers, retries and error handling, but nothing the model reads. Recognise the other rows well enough to spot what a prompt was quietly relying on.

That is exactly Soren's problem. In claude.ai, Anthropic's system prompt gave Claude the date, his project supplied the metric definitions, and the chat window rendered the Markdown. His API request carried none of it. The fix is not a cleverer sentence. It is a system prompt that writes down what the old prompt relied on: a role, today's date, Tallyforge's metric definitions, and an answer format the panel can display.

The staff side is the mirror image. Adaeze, a platform engineer, maintains the team's Claude Code plugin, and a CLAUDE.md file in the repository sets the coding conventions. Both shape her colleagues' Claude Code sessions; neither reaches the explainer's API calls. If the product needs a rule, the product's own request must carry it.

2.5.3 Content boundaries: what gets in, and what gets out

Soren's first request was one long string: his instructions, then the report data, then the user's question. It worked until a customer's report included a campaign named "IGNORE - internal test". The explanation left that campaign out of every total. Nobody attacked anything. The model had no way to tell where Tallyforge's instructions ended and the customer's data began, so it read a campaign name as an order.

A content boundary is the line your application draws around each kind of content: what enters the request, in which place, and what leaves it for the user. On the way in, three kinds of content matter.

  • Instructions. Written by your team, reviewed like code, and placed in the system prompt. They say what the model is for and how to treat everything else.
  • Data. The report, the user's question and any tool results. It changes on every request and often comes from people you do not control, so it travels in the conversation, never in the system prompt. It goes in the user turn inside labelled tags such as <report> and <question>, or in a tool result. Anthropic's prompting guide recommends a tag per kind of content because it reduces misinterpretation.
  • Everything else. Other customers' reports, internal notes, secrets. In a product shared by many customers, each customer is a tenant, and one tenant's data has no place in another tenant's request. The task does not need any of this, so the request never carries it.

Three kinds of content, three decisions

Instructions

System promptwritten and reviewed by your team
Role, date, definitions, format
Says: data is read, not obeyed

Data

Report, question, tool results
Wrapped in labelled tags
Changes on every request

Never sent

Other tenants' data
Secrets and internal notes
Anything the task does not need
The request your code assembles is the boundary: instructions in one place, labelled data in another, and everything else left out.

The boundary runs the other way too. What the model writes goes back through your code before a customer sees it, and your code decides what to display: the fields the panel expects, not raw output. The strongest output boundary is the input one. A model cannot reveal a report that was never in its request. No instruction such as "never mention other customers" is as reliable as leaving them out.

Text written deliberately to hijack a model, known as prompt injection, is a security problem with defences of its own. The design job here comes first: make sure the model can always tell your instructions from the material it works on.

2.5.4 Schema design: contracts for inputs, outputs and tools

The explainer panel shows an arrow badge (up, down or flat) next to each answer. Soren's prompt said "start your answer with the trend". The badge code then met "Up", "increasing", "Sessions rose modestly" and, once, a polite paragraph with no trend at all. Each was reasonable English and a bug in the product.

The fix is a schema: a machine-readable contract for the shape of data crossing a boundary. A Claude application has three of them, and each has its own design rule.

Contract Where it lives Design rule
Input data The user turn, inside labelled tags Send structured data (JSON or a table) with units, the date range and the metric names your system prompt defines; the model cannot guess which week a figure covers or that "sessions" exclude bots
Model output output_config.format with a JSON schema Required fields, an enum for every closed set, and additionalProperties: false so no surprise fields appear
Tool inputs Each tool's input_schema, with strict: true where a malformed call would break something A detailed description of what the tool does and when to use it, and typed, required parameters

Here is the explainer's output contract. Look at the enum on trend, a fixed list of allowed values that turns free text into a closed set, and at additionalProperties, which keeps unexpected fields out of the panel.

explanation_schema = {
    "type": "object",
    "properties": {
        "summary": {"type": "string"},
        "trend": {"type": "string", "enum": ["up", "down", "flat"]},  # a closed set, not prose
        "drivers": {"type": "array", "items": {"type": "string"}},
        "caveat": {"type": "string"},                                  # optional: gaps in the data
    },
    "required": ["summary", "trend", "drivers"],
    "additionalProperties": False,                                    # no surprise fields
}

response = client.messages.create(
    model=MODEL, max_tokens=1024, system=SYSTEM_PROMPT, messages=history,
    output_config={"format": {"type": "json_schema", "schema": explanation_schema}},
)

With structured outputs, the API constrains Claude's reply to this schema, so the badge code receives JSON with one of three trend values. Structured outputs also insists on additionalProperties: false for every object, so that line is a requirement as well as a good habit. The documentation names the exceptions: a refusal or a reply cut off at max_tokens may not match, and the capitalisation of an enum value is not guaranteed. Your code still checks what arrives.

Tool schemas work the same way. A get_report tool with a one-line description and an untyped id invites guesses; a detailed description and a typed, required report_id make the right call the easy one. Setting strict: true on a tool gives its inputs the same guarantee that output_config gives the reply.

2.5.5 Session hygiene: what a conversation carries, and for whom

Now the leak. In testing, a user at the demo shoe shop asked the explainer "Which campaigns have we discussed?", and the answer named a campaign from the demo bakery. The model did not remember the bakery. The Messages API is stateless: each request carries the full conversation, and the model knows only what is in it. If a bakery campaign appeared, Soren's code had put it there.

It had. For the prototype, he stored the conversation in one list at the top of the module. Every request from every customer appended to it, and every request sent all of it. It is a notepad on a shared reception desk: each visitor adds a line and can read every line above. A session is the history your application keeps and resends for one piece of work. Session hygiene is deciding what a session carries, whom it belongs to, and when it ends.

The fix fits in a few lines. Look at the key: each history belongs to one tenant, one user and one conversation, and each request sends only that history. Your server takes the tenant and user from the signed-in session, never from text the browser sends.

histories = {}  # in production: a table keyed the same way, never one shared list

def ask(tenant_id, user_id, conversation_id, question, report_json):
    key = (tenant_id, user_id, conversation_id)    # SCOPE: one tenant, one user, one task
    history = histories.setdefault(key, [])
    turn = f"<question>{question}</question>"
    if not history:                                # a new conversation starts with its own report
        turn = f"<report>{report_json}</report>\n{turn}"
    history.append({"role": "user", "content": turn})
    response = client.messages.create(
        model=MODEL, max_tokens=1024, system=SYSTEM_PROMPT,
        messages=history,                          # this conversation's turns only
    )
    history.append({"role": "assistant", "content": response.content})
    return response

Scoping answers "for whom". Hygiene also answers "for how long". A new report or an unrelated question starts a new conversation, because turns about another report now mislead more than they help. A long conversation about one report gets trimmed or summarised before each call rather than resent whole; doing that well is a context-management skill of its own. And a stored history is customer data, kept only as long as the product needs it.

The same rules apply to the staff's own tools. Each new Claude Code session begins with a fresh context window (everything the model can see at once), and CLAUDE.md files and auto memory are what carry knowledge into it. Within a session, context piles up. When Adaeze moves from one customer's data import to another's, the first customer's error logs and column names are still in context, steering the second job. The /clear command starts fresh between unrelated tasks, and /compact replaces a long history with a summary when she needs to keep going on the same one.

2.5.6 Plugin management: packages that run in every session

Tallyforge's staff side has its own design problem. Adaeze's team plugin is called tallyforge-dev. Over a few months colleagues also installed five more plugins "just in case". Sessions felt heavier, one plugin ran a hook nobody had read, and one Monday the team plugin behaved differently without anyone deciding it should.

A Claude Code plugin is a directory of components that Claude Code installs and loads as one unit; the table below names the main ones. A manifest at .claude-plugin/plugin.json usually names it, and its skills run with that name as a prefix, such as /tallyforge-dev:explain-metric. The same plugin format also installs on claude.ai, where a different set of components loads. None of it reaches a product's API calls.

The part that surprises people: an enabled plugin is part of every session, not only the sessions where you use it.

Component What it adds to every session where the plugin is enabled What to review before enabling it
Skills (instructions Claude loads when relevant) and agents (helpers Claude can hand work to) For those Claude can invoke on its own, the name and description sit in context every turn; the full text loads only when used Whether the team needs them, since idle descriptions still cost context
Hooks Shell commands that run at points such as after every edit, with your user permissions The command each hook runs, in hooks/hooks.json
MCP servers Tool servers that run alongside the session and give Claude their tools Each server's command or URL, in .mcp.json
Manifest The plugin's name, plus an optional version and description Who publishes it, and which marketplace it comes from

Plugin management comes down to three habits. Choose a plugin for a job the team actually does, and disable idle ones: claude plugin disable stops a plugin without uninstalling it. Review what the hooks and servers run before enabling anything, because what a plugin runs, it runs as you. Maintain it like code.

Adaeze installs the team plugin at project scope, which records it in the committed .claude/settings.json so it is enabled for everyone in the repository; each colleague still installs it once. Updates come from the team's marketplace, the catalog that lists the plugin and where to fetch it. Auto-update was on for that marketplace and the manifest set no version, so every commit counted as a new release. That was the Monday surprise: an unreviewed change reached everyone at their next session. Now changes go through review, and the manifest pins a version that Adaeze raises only for a release; until she does, everyone stays on the same copy.

2.5.7 The exam traps

Every trap in this skill comes from forgetting that the model sees only the package. Each wrong practice below looks reasonable on its own; each right one changes what the application assembles.

  • ✗ Assuming a prompt, a CLAUDE.md or a team plugin that works in one interface carries over to your product. ✓ Rebuild what it relied on (role, date, definitions, format, tools) in the product's own system prompt and request. The API sends only what your code puts in it.
  • ✗ Concatenating instructions, data and the user's question into one string. ✓ Put instructions in the system prompt and each kind of data in its own labelled tag, so the model can tell what to follow from what to read.
  • ✗ Telling the model to keep other customers' data secret. ✓ Never send it. The request your code builds is the boundary; an instruction is not a control.
  • ✗ Asking for "JSON, please" and parsing whatever comes back. ✓ Define a schema with required fields, enums for closed sets and additionalProperties: false, and send it with structured outputs or strict tools.
  • ✗ Sharing one conversation history across users, or reusing it across unrelated tasks. ✓ Key every history by tenant, user and conversation, start fresh for new work, and /clear between unrelated tasks in Claude Code.
  • ✗ Installing plugins "just in case" and letting them change without review. ✓ Enable what the team uses, read what its hooks and servers run, and release updates deliberately.

Four tempting fixes, one real one

A bigger modelit sees the same package
Louder wording"NEVER mention other customers"
Config the interface ignoresCLAUDE.md for an API call
Retry until it looks rightthe cause stays
Change what the application assemblescontext, boundaries, schema, scoped session
A bigger model, louder wording, a config file the interface never reads and blind retries all leave the package unchanged; only changing what the application assembles fixes the cause.

2.5.8 Put it together: build a report explainer and break it

You now have every piece: interfaces, boundaries, schemas, sessions and plugins. The quickest way to make them stick is to build a small explainer, then break it twice and watch it fail the way Soren's prototype did.

The artefacts you just made are the start of the rest of the domain. Configuration management (2.6) puts prompts, CLAUDE.md, settings.json and plugin dependencies under version control, with pinned model versions, so a Monday surprise becomes a reviewed release. Context engineering (6.1) shows how to prune and compact the histories you now scope correctly. Output handling (6.3) validates what comes back even when a schema constrains it. AI application security (7.1) hardens the content boundary against text written to attack it.

Key takeaways

  • ✓ The model responds to the package an interface assembles (system prompt, tools, memory and history), so the same prompt can behave differently in claude.ai, Claude Code and the API.
  • ✓ The Messages API sends only what your request carries; rebuild whatever a prompt relied on elsewhere, and remember that CLAUDE.md and plugins shape Claude Code sessions, not your product.
  • ✓ Keep trusted instructions in the system prompt and per-request data in labelled tags, and leave out anything the task does not need, because what is never sent cannot leak.
  • ✓ Design a schema for every boundary: structured input data, a JSON output schema with required fields, enums and additionalProperties: false, and tools with detailed descriptions and typed inputs.
  • ✓ Scope every conversation history to one tenant, one user and one task, start fresh for unrelated work, and use /clear for the same reason in Claude Code.
  • ✓ An enabled plugin is present in every session, so enable only what the team uses, review its hooks and servers, and release its updates deliberately.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

51 CCDV-F questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 51 questions on Applications and Integration alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources