Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 1 › Lesson 1.1

CCAR-P · Domain 1 · 17% of the exam · Lesson 1.1 · 24 min read

Turning a business problem into a Claude solution

How an architect turns a vague request into a scoped Claude solution: frame the problem, test LLM fit, choose app or build, and prove it on real data.

Written against objective 1.1 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

1.1.1 Why "AI for claims" is not a problem yet

Rosalind runs claims at Stonemeadow Mutual, a mid-size property insurer. She comes back from an industry conference with a request for Dariusz, the solution architect: "We need AI for claims. Something like the agent I saw in the demo, one that reads a claim and handles it." The board wants an AI plan from every department by the end of the quarter, and she wants claims to go first. The budget is there. What is missing is a problem.

"AI for claims" names a technology and a department. It does not say what should be different on a Monday morning once it works: faster payments, less retyping, less fraud, happier customers? Each of those outcomes leads to a different system, and some lead to no large language model (LLM) at all. A team that builds from the request as worded will produce something, and it will demo well. Six months later nobody can say whether it helped, because nobody said what helping meant.

Think of a patient who walks into a clinic asking for a drug they saw advertised. A good doctor neither writes the prescription nor refuses it. They ask about the symptoms first, because the drug is only right if it treats what is actually wrong. Translating a business problem into a Claude solution is the architect's version of that consultation. You frame the problem in business terms, split it into jobs and give each job the tool that does it best. Then you choose how the Claude part is delivered and state what success looks like before anything is built.

Two ways to answer "AI for claims"

Technology first

"We need an agent"
Build a demoon invented claims
Look for a problem it solves
Impressive, unmeasurable

Problem first

The outcomeand today's baseline
Split the request into jobs
The right tool for each job
Prove it on real claims
Starting from the technology produces a demo in search of a problem; starting from the problem produces a scoped design with a test it can pass or fail.

1.1.2 Frame the problem before you design anything

Here is the question that separates architects from enthusiasts: what would you need to know before you could say whether ANY design is good? Five facts, and none of them is about Claude. Together they are the problem frame, and every later decision (the tool, the delivery form, the success criteria) is justified by pointing back at one of them.

Question Why it decides the design Stonemeadow's answer
OUTCOME: what should be different, in business terms? It becomes the success criterion; "use AI" is not an outcome Simple claims get first contact the same day, not after three working days
WORK TODAY: who does it, and where does it stall? Shows which steps already work and where the time goes 38 handlers read each new claim and retype it: about 20 minutes a claim
VOLUME: how many, how fast, how spiky? Sets cost and latency, and whether a person starts each task About 4,000 new claims a month, tripling in the week after a storm
ERROR COST: what does each kind of mistake cost? Decides where a person must confirm and how cautious the system must be Fast-tracking the wrong claim can pay out on an inflated or uncovered loss; missing a fast-track keeps today's delay
DATA: where does each input live, and who owns it? Decides integrations, access and whether the data may be used at all Reports, call notes and photos in a document store; policies in the policy system; a fraud score in the claims system

Two rows deserve a second look. The error-cost row is usually asymmetric, and the asymmetry shapes the design more than any accuracy figure. At Stonemeadow, fast-tracking the wrong claim costs money and invites fraud, while missing a fast-track leaves a customer where they are today. That single fact already says the system should flag and a handler should confirm, and that the flag should lean cautious.

The work-today row shows where the time really goes. The bottleneck is the first notice of loss (FNOL), the customer's first report of a claim, arriving as a web form, an email or a call-centre note, often with photos. Handlers are not slow at deciding; they are slow at reading free text and retyping it. Anthropic's ticket-routing guide makes the same point before any build: study how people handle the work today, including which automated rules already exist and where they fail. Drawing these answers out of busy people is a skill of its own; what matters here is having them before you design.

With the frame in hand, Dariusz rewrites the request as an outcome Rosalind can sign. Simple property claims get first contact the same day instead of after three working days, with no more money paid on claims that should never have been fast-tracked.

1.1.3 Decide whether a language model is the right tool

Here is the mistake an AI budget invites: once the money says "AI", every part of the request looks like an AI job. Resist it. A business request is usually several jobs bundled together, and each has its own best tool. Unpacked, Rosalind's "AI for claims" is four jobs. Read each new claim so nobody retypes it. Spot the simple ones for fast-track. Check that the policy covers the loss and work out the excess (the deductible). And catch fraud, "including doctored photos".

Tool When it wins What it costs
Rules in code The logic can be written down exactly, and the answer must be identical and auditable every time Brittle on free text; every new case is a code change
Existing software A system of record already does the job Nothing to build, but its limits stay
Classic machine learning (ML) Years of labelled, structured history, a stable target, a score needed at volume Large labelled datasets, retraining when categories change, little explanation
Claude Unstructured text or images, rules that depend on meaning, few labelled examples, changing categories, reasons wanted Probabilistic output, a cost and a wait per call, and tests on real cases plus human checks where errors are expensive

Most of the Claude row comes from Anthropic's ticket-routing guide. Its signs that favour an LLM over traditional machine learning also include ambiguous edge cases and several languages, and it notes that Claude can classify from a few dozen labelled examples. The mirror image matters just as much. When the answer must come out the same every time and trace back to a rule, or a system already owns it, a model adds cost and variability and buys nothing. Claude also knows nothing about Stonemeadow's policies or customers except what your code puts in the request, and it can write fluent, confident text that is wrong. That is tolerable in a summary a handler checks, not in a coverage decision.

You would not ask your most articulate colleague to run the payroll by hand, and you would not ask the payroll system to read a customer's angry letter. Two of Rosalind's four jobs leave Claude's list for exactly that reason. Checking cover (was the policy in force on the date of loss, is the peril insured, what is the excess) is exact, auditable logic the policy system already runs. Dariusz keeps it there and feeds it the extracted fields. Fraud stays with the existing fraud model, trained on years of confirmed outcomes.

The "doctored photos" part of the fraud job is ruled out by a documented limit, not by cost. Anthropic's vision documentation says Claude cannot determine whether an image is AI-generated and should not be relied on to detect fake images. Claude also receives no image metadata, so checking when a photo was taken is a job for code.

One request, four jobs, three kinds of tool

Rules in the policy system

Cover and excessexact, auditable, already built

Fraud model and code

Fraud scoretrained on years of outcomes
Photo datesread from metadata by code

Claude

Read the report and photossummary and fields
Spot simple claimsa flag with reasons
Only the jobs that need reading and judgment go to Claude; exact logic, the fraud score and photo metadata stay with tools built for them.

1.1.4 Map each job to a Claude capability

Once a job belongs to Claude, name the capability precisely, because each one implies a different prompt, a different output and a different test. "Claude will handle intake" cannot be evaluated; "Claude extracts eleven fields into a fixed schema" can. Six capabilities cover most business requests.

Capability What it produces At Stonemeadow
Generation New text: a draft, a reply, a letter Not in the first release; a draft first-contact message is a later candidate
Extraction Fields from unstructured input, in a fixed structure Policy number, date of loss, address, peril and rooms affected, for the claims system
Classification A label from a defined set, with the reason "Likely fast-track" or "needs a handler", with the evidence
Summarisation A shorter account for a named reader Five lines for the handler: the report, the call notes, what the photos show
Reasoning over documents A judgment that connects several sources Does the damage in the photos match the story in the report?
Tool use and agents Requests for actions your code runs, possibly in a loop Not needed: the steps are known in advance, so code runs them

Two rows carry most of the design. Extraction feeds a system of record, so its shape must be exact. The Claude API's structured outputs feature constrains a reply to a JSON schema you supply, so the fields can land in typed columns. The docs list the exceptions, such as a reply cut off by the token limit. A valid shape is not a true value, though: a well-formed policy number can still be the wrong one, so the values still need testing.

The last row is the one people reach for first. Tool use lets Claude ask your code to run a function, such as a policy lookup; an agent lets Claude choose its own next step in a loop until the task is done. Tool use earns its place when Claude must decide what to fetch or do, and an agent when even the number and order of steps depend on what each step finds.

Stonemeadow's intake runs the same steps in the same order every time, so plain code calls Claude, then the policy system, then the fraud model. Anthropic's guidance on building agents says the same: find the simplest solution possible and add complexity only when needed, which may mean not building an agent at all.

The intake assistant, job by job

NEW CLAIMreport, call notes, photos
CLAUDEsummary, fields, flag with reasons
YOUR CODEcover check, fraud score
HANDLERconfirms or overrides the flag
Claude reads and judges; your code gets cover from the policy system and a score from the fraud model; a handler confirms every fast-track before anything is paid.

1.1.5 Choose the delivery form: a Claude app or a custom build

Now a question that is easy to skip: does this need building at all? Many requests are met by the Claude apps on a business plan. A Team or Enterprise plan gives each person Claude on web, desktop and mobile, plus Projects that keep a team's instructions and reference files together. It adds connectors that search and retrieve from tools such as Microsoft 365, Google Drive and Slack. Enterprise adds controls a security team asks for, such as audit logs and custom data retention. A custom build means your own software calls Claude through the Claude API or a cloud platform such as Amazon Bedrock, Google Cloud or Microsoft Foundry.

Two requirements decide between them: who starts the work, and where the result must land. If a person starts each task, the work varies and the output is something people read, the app wins. If the work must run on every case without anyone starting it, write a fixed shape into a system of record, or apply one tested prompt to every case, you are building.

Delivery form When it wins What it costs
Claude app on a Team or Enterprise plan, with Projects and connectors A person starts each task; the work varies; the output is read, not stored; the data sits where a connector reaches Nothing runs unless someone asks; no fixed schema into your systems; consistency depends on each user's prompting
Custom build on the Claude API Every case, automatically; structured output into your systems; one tested prompt and schema You build, evaluate, monitor and run it
Custom build through a cloud platform As for the API, and your data, identity and billing already live in that cloud Feature availability differs by platform, so check each one the design needs

The intake assistant has to run on all 4,000 claims a month, before any handler opens them, and write fields into the claims system. That is a custom build. Stonemeadow's claims platform, identity and billing already run in AWS. Its security team adds a sharper requirement: claim data may go only to a model service that AWS itself operates, inside the AWS boundary the team already audits. That rules out Claude Platform on AWS, which Anthropic operates behind AWS sign-in and billing, and points to Claude in Amazon Bedrock, which AWS runs.

Dariusz then checks every feature the design uses against the docs for that platform, and the check pays off. Bedrock offers structured outputs today only on some older Claude models, not on the newest Opus and Sonnet, while the Anthropic-operated route has them. The security requirement wins, so unless the proof of concept picks one of those older models, the claims system's code must validate every reply against the schema before it writes a field. Dariusz writes the trade-off down, so the security team sees what its requirement costs.

The app still has a place. Senior handlers who want help reading an engineer's report or drafting a difficult letter start that work themselves, and each case is different. Once the security team has reviewed the plan's controls, they get Enterprise seats and a shared Project holding the claims-handling guide. One engagement, two delivery forms, each justified by its own requirement.

1.1.6 Turn the design into a hypothesis you can prove

Rosalind's next question is the one every sponsor asks: "Will it work?" The honest answer is a solution hypothesis. It says what Claude will do and for whom, which business number should move from what baseline to what target, what is out of scope, and what result would prove you wrong. It turns a promise into a test.

Write its success criteria in business terms, against today's baseline. "95% accuracy" means nothing to Rosalind; "simple claims contacted the same day" does. Anthropic's guide to defining success asks for criteria that are specific, measurable, achievable and relevant, and notes that most use cases need several at once. Achievable means anchored in evidence. Stonemeadow's handlers already get about 3 fields in 100 wrong when retyping, so 97% field accuracy is parity with people: a defensible bar, not a hopeful one.

Here is the hypothesis Dariusz takes to Rosalind. Look at the out-of-scope line, which records the jobs ruled out, and the last line, which says in advance what result would stop or reshape the project.

Problem: simple property claims wait an average of three working days for first contact, and handlers spend about 20 minutes retyping each first notice of loss.
Hypothesis: if Claude summarises each new claim and its photos, extracts the policy and incident fields, and flags likely fast-track claims with its reasons, handlers can contact simple claims the same day.
Out of scope: cover and the excess (rules in the policy system); fraud scoring and photo-metadata checks (the existing fraud model and code); any payment or customer message without a handler.
Success, against today's baseline: at least 97% of extracted fields match the corrected claim record (handlers today: about 97%); fewer than 1 in 50 flagged claims prove unsuitable for fast-track; same-day first contact on flagged claims during the pilot.
Proof of concept: 300 closed claims from the last twelve months across web, email and phone, including blurred photos, missing policy numbers and reports in other languages; the truth comes from the settled claim file.
Stop or reshape if: field accuracy stays under 90% after two prompt revisions, or more than 1 in 20 flagged claims prove unsuitable.

Then scope the proof of concept (PoC) to the riskiest assumption, not the whole product. Here that is whether Claude can read Stonemeadow's messy reports and photos accurately enough. Run it on real samples. Closed claims are ideal, once the data owner has agreed to their use, because their outcome is already known: the corrected fields and whether the claim settled cleanly. Anthropic's eval guidance asks for tests that mirror your real task distribution, edge cases included, which is why the 300 span every channel and peril. The PoC tests the first two criteria offline; the same-day criterion needs a short pilot with real handlers.

A proof of concept that can fail

SAMPLE300 closed claims, outcomes known
RUNone prompt, one schema
SCOREagainst criteria set in advance
DECIDEbuild, reshape or stop
The criteria and the stop line are fixed before the run, so the result decides the next step instead of the demo.

1.1.7 The exam traps

Every trap here skips a step of the translation: it starts from a tool, gives Claude a job it should not have, or declares success without a test.

  • ✗ Starting from the technology: "we need an agent", "put the biggest model on it". ✓ Start from the outcome and today's baseline. The pattern and the model follow from the jobs, and the simplest design that meets the requirement wins.
  • ✗ Giving every part of the request to Claude. ✓ Split it into jobs and keep exact, auditable logic in rules and existing systems. A coverage check does not improve by becoming probabilistic.
  • ✗ Asking Claude for something its documented limits rule out, such as spotting AI-generated photos. ✓ Check the docs' limits during design and route those jobs to tools built for them, such as the fraud model or code that reads metadata.
  • ✗ Defaulting to a custom build, or to the apps, before asking who starts the work. ✓ Decide by who starts the work and where the output lands. Person-started, varied work fits the apps; every-case work into a system of record needs a build.
  • ✗ Success criteria in model terms only ("95% accuracy"), with no baseline. ✓ State them in business terms against today's numbers, with the model measures that feed them and a stop condition.
  • ✗ Proving the idea with a demo on invented or hand-picked examples. ✓ Run a proof of concept on real, representative samples with known outcomes, edge cases included, scored against criteria set beforehand.

1.1.8 Put it together: turn a request into a scoped proof of concept

You now have the whole translation, from a request worded as technology to a hypothesis that a real sample can pass or fail. Stonemeadow's "AI for claims" has become an FNOL intake assistant that summarises, extracts and flags, with two jobs deliberately left where they already work. The fastest way to own the method is to run it once, small.

This lesson ends where design begins. End-to-end architecture (1.2) turns the intake assistant into stages with defined inputs, outputs and a feedback loop. Pattern selection (1.3) is the full version of "code runs the fixed steps": when a workflow, an augmented model call or an agent fits. Business value pillars (1.6) turn the success criteria into key performance indicators and a return-on-investment case the board can weigh. Structured discovery (6.1) is how you draw the five framing answers out of people who each know only part of the process.

Key takeaways

  • ✓ A request that names a technology ("AI for claims", "an agent") is not a problem yet; start from the outcome the business wants and today's baseline.
  • ✓ Frame every problem with five facts: the outcome, who does the work today, the volumes, what each error costs, and where the data lives.
  • ✓ Split the request into jobs: rules and existing systems for exact, auditable logic, classic ML for scoring on labelled history, and Claude for unstructured language, images and judgment.
  • ✓ Name each Claude job as a capability such as extraction or classification, and add tool use or an agent only when Claude must decide what to fetch or which step comes next.
  • ✓ Choose the delivery form by who starts the work and where the output lands: a Claude app on a Team or Enterprise plan, or a custom build on the API or a cloud platform, checked feature by feature.
  • ✓ Write a solution hypothesis with business success criteria, a baseline, an out-of-scope list and a stop condition, and prove it on real samples with known outcomes.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

33 CCAR-P questions on Domain 1, free

Every question in the bank is tagged to a domain, so you can drill 33 questions on Solution Design & Architecture alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 1 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources