Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 2 › Lesson 2.1

CCDV-F · Domain 2 · 33.1% of the exam · Lesson 2.1 · 21 min read

Turning business needs into functional and infrastructure requirements

How to turn business wishes into testable functional and infrastructure requirements for a Claude app, trace them to the architecture and raise open questions.

Written against skill 2.1 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

2.1.1 Why "fast, safe and cheap" is not a specification

Picture the kickoff meeting. Henrike runs operations for a property-management company with about 30,000 rented homes, two thirds in Germany and the Netherlands, the rest in the United States. Every winter her office team drowns in tenant messages. Where is my deposit? When is the boiler engineer coming? Can I keep a cat? She wants an assistant that answers tenants by chat and email, and she describes it in one breath. "It has to be fast, it must never give legal advice, European data stays in Europe, and it can't cost more than the temps we hire every January."

Kwame, the developer, could open his editor that afternoon. A call to Claude that answers "can I keep a cat?" takes ten minutes to write, and the demo would look impressive. It would also be built on guesses. "Fast" might mean two seconds or two hours. "Europe" might mean where the model runs, where transcripts are stored, or both. Nobody has said how many messages arrive on the Monday after a storm. Each guess becomes a design decision, made silently by whoever writes the code first.

The work that comes before the code is requirements: turning what the business wants into statements a developer can build to and a tester can check. For a Claude application they come in two kinds, functional (what the feature does) and infrastructure (the conditions it runs within). Each traces back to a business reason and forward to a part of the architecture, and whatever nobody has decided becomes an open question.

From a wish to something you can build and test

What Henrike said

"It has to be fast"
"Never give legal advice"
"European data stays in Europe"
"No more than the winter temps"

What Kwame writes down

Chat: first words within 2 secondsemail: a reply within 15 minutes
Every legal question goes to a personproven on a test set
EU messages processed on an EU endpointlogs and storage: open question
Under 8,000 euros a month at forecast volumealert at 80%
Each phrase from the kickoff meeting becomes a requirement with a number, a home in the architecture and a test.

2.1.2 Three layers: business, functional, infrastructure

Here is the question that trips people up: when Henrike says "European data stays in Europe", is that a feature? It is not. Nothing the tenant sees changes. It is a condition on where and how the feature runs, and mixing the two up is the root of many bad designs. So the first skill is sorting every statement into the right layer.

A business requirement is the outcome the organisation wants, in its own words: cut the reply time from two days to minutes, take routine questions off the office team, stay out of legal trouble. It says WHY, and rarely anything a program can check.

A functional requirement says what the Claude feature must do to serve that outcome, in four parts. The inputs: the tenant's message, language, lease and open repair tickets. The outputs: a reply in the tenant's language, short for chat and a full letter for email. The behaviour at the boundaries: a legal question, a gas leak, a question the records cannot answer. And the acceptance criteria, the measurable tests that say the requirement is met.

An infrastructure requirement sets a condition the feature must run within: speed, peak volume, reliability, region, how long data is kept, who may see what, cost and provider. Tenants never read these, but break one and the project fails anyway. Think of renovating a kitchen in one of Henrike's flats. "A family needs to cook here" is the business need, and four hobs and a dishwasher are the functions. Wiring that carries the oven's load, the building code and the budget are the infrastructure. A beautiful kitchen that trips the fuse box is still a failed renovation.

Layer The question it answers In the tenant assistant
Business Why are we building this? "Tenants wait two days for an email reply"
Functional What must the feature do, and how do we test it? Answer lease, rent and repair questions from the tenant's own records; at least 90% correct on a test set
Infrastructure Under which conditions must it run? EU tenants' messages processed in the EU; first chat words within 2 seconds; under 8,000 euros a month

Memorise the three questions in the middle column; they let you sort any statement in a scenario into a reason, a behaviour or a condition.

2.1.3 Functional requirements you can test

It is tempting to write "the assistant answers tenant questions accurately" and move on. Resist it. "Accurately" gives nobody a line to measure against, and a model's output varies with the wording of a question and from one run to the next, so one good demo proves little. A requirement that cannot fail cannot pass either. Anthropic's guidance asks for success criteria that are specific, measurable, achievable and relevant. Its own example sets a bad criterion, "safe outputs", beside a good one: fewer than 0.1% of 10,000 outputs flagged for toxicity.

So an acceptance criterion for a Claude feature is usually a rate on a test set, a fixed collection of realistic inputs with known right answers. Kwame builds his from last winter's tenant emails, stripped of names. FR-1 reads: "On 300 past questions, at least 90% of replies are judged correct against the lease and house rules." Some behaviours tolerate no misses. FR-2: "All 40 legal-dispute and emergency messages go to a person." Kwame proposes the thresholds, but Henrike signs them off, because how many errors tenants may see is a business decision. The set deliberately includes awkward cases: Dutch and German messages, questions the records cannot answer, a furious tenant.

The most important functional requirements are often about what the feature must NOT do, and what it does when it does not know. FR-3: when the answer is not in the tenant's records, it says so and opens a ticket for a person. It never promises a repair date the ticket system does not hold. Each rule is testable, so each gets its own cases.

Here is an acceptance criterion once it becomes code. Look at the CASES list, which is the requirement's test set, and at the last line, which compares the pass count with the target. MODEL and TRIAGE_PROMPT stand for your own model choice and routing instructions.

import anthropic

client = anthropic.Anthropic()
CASES = [  # FR-2 test set: past messages, each with the route the requirement demands
    ("My landlord wants to evict me. Is that legal?", "HUMAN"),
    ("Ich rieche Gas im Treppenhaus", "EMERGENCY"),   # German: "I smell gas in the stairwell"
    ("When is my rent due this month?", "ANSWER"),    # routine: must NOT be escalated
]
passed = 0
for message, expected in CASES:
    reply = client.messages.create(
        model=MODEL, max_tokens=1024, system=TRIAGE_PROMPT,
        messages=[{"role": "user", "content": message}],
    )
    label = "".join(b.text for b in reply.content if b.type == "text").strip()
    passed += label == expected                       # exact match: one right route per case
print(f"FR-2: {passed}/{len(CASES)} routed correctly (target: all of them)")

The third case keeps the test honest: a triage that sent every message to a person would pass the first two and fail it. A real test set has dozens of cases, and some criteria, such as "correct against the lease", need a grader smarter than an exact match. The requirement's job is to name the test and the target; how you grade comes later.

2.1.4 Infrastructure requirements: the conditions it runs within

Henrike's "fast" hides the next trap: one application often contains several workloads with different needs. A tenant in the chat window is watching the screen, so what matters is time to first token, the delay before the first words appear. Tokens are the word pieces the model reads and writes, and the unit the API counts. A tenant who sent an email is not watching, so 15 minutes is generous. One figure for both would be too strict for email or too loose for chat. So Kwame writes requirements per workload, numbered IR-1 to IR-8, and states latency as a p95: 95 replies out of 100 must be at least this fast.

Requirement The question to put to the business Tenant assistant (draft)
IR-1 Latency How long may a user wait, per channel? Chat: first words within 2 s at p95; email: a reply within 15 minutes
IR-2 Throughput What is the peak volume, not the average? 2,500 messages a day; 25 in the worst minute of last winter
IR-3 Availability What must happen when a call fails? Chat promises an email reply; email waits and retries; no message is lost
IR-4 Residency Where must data be processed and stored? EU tenants in the EU, US tenants in the US
IR-5 Retention How long may each copy of the data live? Transcripts 90 days (to confirm); logs without message text
IR-6 Security Who may see what, and who may act? The tenant's identity comes from the login, never from the message
IR-7 Cost ceiling What is the budget, at which volume, and what happens at the limit? Under 8,000 euros a month at forecast volume; alert at 80%
IR-8 Provider Which provider and cloud, under which contracts? Decided by residency (see the next section)

Throughput needs the most care, because the Claude API measures capacity in its own units. Each organisation has limits per model in requests, input tokens and output tokens per minute, set by its usage tier, a level Anthropic assigns from usage history. Exceed any one and the API returns HTTP status 429, a rate-limit error, with a retry-after header saying how long to wait. A per-minute limit may also be enforced over shorter intervals, so a burst inside one minute can trip it. A daily average tells you almost nothing.

So Kwame sizes from the worst minute in last winter's logs: 25 messages, each needing about three model calls of roughly 4,000 input and 300 output tokens. That makes 75 requests, 300,000 input tokens and 22,500 output tokens a minute, to compare with the limits of whichever provider serves the traffic. Published limits are ceilings, not guaranteed capacity, and a sharp jump in usage can trigger separate acceleration limits, so the launch ramps up gradually.

Two more rows need a note. Availability is a requirement on YOUR design: even a healthy API can return an overloaded error when traffic is high for everyone, so the requirement says what the tenant sees then. Retention covers every copy of the data: your database, your logs and the provider. Anthropic offers zero data retention (ZDR) for eligible API features, agreed per organisation: under it, Anthropic stores no prompts or responses once the response is returned. On Amazon Bedrock and Google Cloud, the cloud provider handles the data under its own terms.

The cost ceiling needs a volume and a consequence. "Under 8,000 euros a month" holds only at the forecast traffic, so the requirement states both. It must also say what happens at the limit. The Claude API, for example, lets an organisation set a monthly spend limit, after which it refuses requests. A hard stop like that turns a budget rule into an outage, so Henrike chooses an alert at 80% and a review instead.

2.1.5 Tracing every requirement to the architecture

A list of requirements is not yet a design. What connects them is traceability: every requirement points back to the business need it serves and forward to the component that satisfies it and the test that proves it. Check both directions. A requirement with no component is a promise nobody will build; a component with no requirement is cost with no reason.

The solution architecture is the set of parts and how they connect. Kwame's draft has seven, and he labels each with the requirements it satisfies.

The tenant assistant, with its requirements attached

Channelschat widget, email inbox: IR-1, IR-3
Gatewaylogs the tenant in, paces traffic: IR-2, IR-6
Triageanswer, person or emergency: FR-2
Claude, EU or US endpointpicked by the tenant's region: IR-4, IR-7, IR-8
Tenant-scoped toolslease, repair tickets: FR-1, IR-6
Handoff queuea person takes over: FR-2, FR-3
Logs and transcriptsin-region, deleted on schedule: IR-4, IR-5
Every box carries at least one requirement, and every requirement lands on at least one box; a gap in either direction is a design problem.

Some requirements do more than land in the architecture: they choose it. Residency is the clearest case. On the Claude API, the inference_geo parameter sets where inference runs, that is, where the model processes each request. At the time of writing it accepts only "global" (the default) and "us", and the only location for data the API stores at rest is the US. So "EU messages are processed in the EU" points to a cloud provider: Amazon Bedrock has endpoints in several EU regions, and Google Cloud an eu multi-region endpoint. Kwame's draft sends both regions through one provider, each to its own regional endpoint.

That choice has consequences. These regional endpoints cost 10% more than the provider's global ones. Not every model is offered in every region, so the later model shortlist may shrink. The provider's quotas replace the Claude API's tier limits in the throughput check, and its data terms replace Anthropic's in the retention check. One sentence from the kickoff meeting has picked the provider, moved the cost estimate and narrowed the models. That is why requirements come first.

Security traces the same way. "A tenant never sees another tenant's data" cannot trace to the prompt, because an instruction to Claude is guidance, not enforcement. It traces to the gateway, which takes identity from the login, and to the tools, which return only that tenant's records.

2.1.6 Assumptions and open questions

Every requirements draft is full of things nobody said. "Tenants write in English." "The 2,500 a day are spread evenly." "Europe means the model call." Each is an assumption, something you treat as true without confirmation. The dangerous ones are those you do not notice you made, because they turn into architecture without anyone deciding.

Surfacing them is a habit. Whenever you write a number or rule that nobody gave you, mark it as an assumption. Whenever a requirement depends on something you cannot decide, such as a legal reading, a budget or a contract, record an open question with the person who can answer it and a date. Some you close yourself with data: Kwame reads the tenant database for languages and last winter's email logs for the peak minute. The rest go to their owners.

Four silent guesses, one habit

"Everyone writes in English"Dutch and German tenants
"2,500 a day, spread evenly"25 in the worst minute
"Europe means the model call"and the logs and backups?
"Keep transcripts for debugging"for how long, and who decides?
Write it down, name an owner, decide before buildingresidency and retention: data protection officer, by 15 October
Each guess would have become a design decision nobody made on purpose; written down with an owner, it becomes a question that gets answered before the build.

The residency question shows why this matters. Kwame's first draft assumed that "Europe" meant only where Claude runs. Asked directly, the company's data protection officer says transcripts and logs must stay in the EU too, and must be deleted after 90 days. That changes the storage design and the logging setup. Learning it in the first week costs an email; learning it in a launch review costs a rebuild.

2.1.7 The exam traps

Most mistakes here skip the translation step: building straight from the wish, or letting one number, one workload or one prompt stand in for a requirement. Each has a precise correction.

  • ✗ Starting from the model or the prompt ("use the most capable model and tune later"). ✓ Derive measurable functional and infrastructure requirements first; the model is chosen afterwards, to meet them at the lowest acceptable cost.
  • ✗ Accepting "fast", "accurate" or "secure" as requirements. ✓ Give each a threshold and a test: first chat words within 2 seconds at p95, at least 90% correct on 300 real questions.
  • ✗ One set of numbers for every workload. ✓ Write requirements per workload: live chat and email replies need different latency targets.
  • ✗ Putting a control into the prompt ("tell Claude not to show other tenants' data"). ✓ Trace it to the architecture: identity from the login, tenant-scoped tools, a region-bound endpoint. A prompt guides; the system enforces.
  • ✗ Sizing throughput from the daily average. ✓ Size it from the peak, converted to requests and tokens per minute, and compare it with the rate limits before launch.
  • ✗ Filling gaps with silent assumptions. ✓ Record assumptions and open questions with an owner and a deadline, and settle them before the design depends on them.

2.1.8 Put it together: write a requirements pack and test it

You now have every piece: three layers, functional requirements with acceptance criteria that can fail, infrastructure requirements per workload, traceability to the architecture, and a habit for assumptions. The fastest way to make it stick is to write a small pack for a real feature, turn one requirement into a test, and watch that test catch a regression.

The rest of Domain 2 builds on this pack. The systems life cycle (2.2) carries these requirements through build, release, operation and maintenance. Claude API mechanics (2.3) turn latency and volume requirements into concrete choices such as streaming and batch processing. And in Domain 5, model selection (5.3) picks the tier that meets the quality, latency and cost targets you wrote here.

Key takeaways

  • ✓ Business requirements say why; functional requirements say what the Claude feature does; infrastructure requirements say the conditions it must run within.
  • ✓ A functional requirement names inputs, outputs, boundary behaviour and an acceptance criterion that can fail, usually a measured rate on a representative test set.
  • ✓ Infrastructure requirements cover latency, throughput, availability, residency, retention, security, cost and provider, and they are written per workload.
  • ✓ Size throughput from the peak in requests and tokens per minute, and compare it with the rate limits that apply before launch.
  • ✓ Trace every requirement to its business reason, an architecture component and a test; residency can decide the provider, and controls belong in code, not in prompts.
  • ✓ Record assumptions and open questions with an owner and a deadline, and settle them before the design depends on them.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

51 CCDV-F questions on Domain 2, free

Every question in the bank is tagged to a domain, so you can drill 51 questions on Applications and Integration alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 2 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources