Home › Study guides › CCAR-P › Domain 1 › Lesson 1.3
CCAR-P · Domain 1 · 17% of the exam · Lesson 1.3 · 25 min read
Augmented LLM, workflow or agent: choosing the pattern
How to choose between one augmented LLM call, a workflow and an agent: the five workflow patterns, when autonomy pays off, and what frameworks hide.
Written against objective 1.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
1.3.1 One inbox, three kinds of request
Keelhaven Freight is a freight forwarder: it books space with shipping lines, airlines and hauliers for companies that need goods moved. About 1,500 quote requests reach its inbox every week, and a pricing desk of eight people answers them by hand, with a median reply time of five working hours. Shippers often book with whoever answers first. So Bruno, the commercial director, has asked for "an AI agent on the quote inbox", and Anjali, the solution architect, has to design it.
First, Anjali reads 400 past requests and sorts them. About 70% are standard lanes: a regular route with a price already in the rate card, such as one 40-foot container from Rotterdam to Singapore. About 20% are special cargo (dangerous goods, refrigerated loads, oversized pieces) that need extra checks and surcharges. The last 10% are unusual: multi-leg project cargo, consignees (the receiving companies) that need screening against sanctions lists, requests with half the details missing.
A plain model call handles none of these well. It has never seen Keelhaven's rate card, so it would invent a plausible price, which is worse than no reply. Something has to be built around it, in one of three broad shapes. One model call can be given the right data and tools, a design called an augmented LLM (large language model). A workflow runs model calls along paths your code fixes in advance. An agent lets the model choose its own next step. The request said "agent". The inbox suggests each kind of request may deserve a different design.
What Keelhaven's inbox actually contains
Standard lane about 70%
Special cargo about 20%
Unusual about 10%
1.3.2 The augmented LLM: the building block
Start with the question an architect asks before any pattern: what is the smallest design that could answer a standard-lane request correctly? The model can already read a messy email and write a clear reply. What it lacks is Keelhaven's data. Give it that data, and one call may be all the job needs.
That design is the augmented LLM: a single model call enhanced with three things. Retrieval places relevant data in its context, tools are functions it can ask your code to run, and memory is a store of notes from earlier work that the model can read and add to. Anthropic's guidance on building agents calls it the basic building block of agentic systems. Every step of a workflow and every turn of an agent is one of these calls. Current models use the augmentations actively: they write their own search queries, pick their tools and decide what is worth keeping.
The augmented LLM for a standard lane
get_surcharges, when the date changes the priceFor standard lanes, Anjali's design is exactly this. Your code finds the ports named in the email, fetches the matching rate-card rows and places them in the prompt beside the email. The model extracts the shipment details, applies the right row and drafts the reply, calling get_surcharges only when the sailing date changes the price.
Notice the choice inside that design: retrieval can be pushed or pulled. Your code can fetch data before the call, or the model can fetch it with a tool when it decides it needs it, the "just in time" approach in Anthropic's context-engineering guidance. Pulling finds what you did not anticipate, but runtime exploration is slower than handing over prepared data, and a hybrid can push some data up front and let the model pull the rest. Here the email names the lane, so pushing the rows saves a round trip and ties every price to the rows the model saw.
Anthropic's guidance is blunt: for many applications, optimising a single call with retrieval and in-context examples is usually enough. Get two things right: tailor each augmentation to the use case, and give the model an easy, well-documented interface to it.
1.3.3 Workflows: your code lays the track
Special cargo is where the single call cracked. In testing, it quoted a pallet of lithium batteries at the plain rate, never asking for the UN number that identifies a dangerous good. Stretching the prompt to cover dangerous goods then made its standard-lane answers worse, a known effect: optimising one prompt for one kind of input can hurt it on others. The fix is not a longer prompt but more than one step.
A workflow is a system in which model calls and tools run along code paths you define in advance. Think of a railway. The model drives each train, but your code laid the track and throws every switch, so a train can only reach stations you built. Anthropic describes five workflow patterns that recur in production, and each fits one shape of task.
| Pattern | What your code does | When it fits, at Keelhaven |
|---|---|---|
| Prompt chaining | Runs fixed steps in order, each call working on the previous output, with checks (gates) between them | The task splits cleanly into fixed steps, and you accept more latency for accuracy: extract cargo details, check them, price, draft |
| Routing | Classifies the input, then sends it down a specialised path with its own prompt, tools or model | Distinct categories are better handled apart and can be classified accurately: standard lane, special cargo, unusual |
| Parallelisation: sectioning | Splits the task into independent parts that run at the same time, then combines them | The parts do not depend on each other: price the sea leg and the trucking leg at once |
| Parallelisation: voting | Runs the same task several times and combines the answers | One attempt is not confident enough: three independent checks of "is this dangerous goods?", flagged if any says yes |
| Orchestrator-workers | A central model call decides the subtasks at run time, hands them to worker calls and combines the results | You cannot predict the subtasks: a project shipment whose number of legs varies |
| Evaluator-optimizer | One call drafts, another evaluates the draft against criteria and feeds back, and the loop repeats until it passes | Criteria are clear and refinement measurably helps: every quote states its validity date, trade terms and surcharges |
Memorise the five and the task shape each fits. Watch orchestrator-workers. It looks like sectioning, but in sectioning your code fixes the parts in advance, while an orchestrator decides them for each input. That puts it one step short of an agent, since your code still fixes the path. Once the workers become agents themselves, you are designing a multi-agent system.
Anjali's design, shown below, puts routing at the front. Two details matter. The standard path can run on a smaller, faster model, because routing lets each path choose its own. And the special-cargo gate is code, not a prompt: it refuses to price until the required fields are present, such as a UN number, a set temperature, or dimensions and weight.
Keelhaven's routing workflow
Standard lane
Special cargo
Unusual
the branch still waiting for a design
1.3.4 Agents: the model chooses the route
That leaves the unusual 10%. Take one realistic request: a 38-tonne transformer from Antwerp to a mine high in the Andes, needed on site by March. Which Chilean port can lift 38 tonnes, which sailings reach it in time, can the mountain road take the load, does the consignee pass screening? Each answer changes the next question. A workflow would need a branch for every answer, and the next odd request would still find a gap.
An agent is a model directing its own tool use in a loop: it decides which tool to call, your code runs it, the result comes back, and the model decides again. Anthropic's context-engineering guidance puts it as "LLMs autonomously using tools in a loop". If a workflow is a railway, an agent is a four-wheel drive given a destination, picking the road at every junction from what it sees. Your code still holds the keys: which tools exist, what each can touch and when the trip ends.
Two features make that workable in production. With environment feedback, the agent gets ground truth from a tool result at each step, such as the port's real crane limit or the actual sailing list. It plans from that, not from its own assumptions. With stop conditions, the loop ends when the model is finished, but you also design the other exits, including where it must pause for a person.
The agent loop and the exits you design
| Exit | What triggers it | Keelhaven's setting |
|---|---|---|
| Finished | The model stops asking for tools | Save the draft quote for the pricing desk |
| Blocker | Information only the shipper has | Draft a question to the shipper instead of guessing |
| Checkpoint | A step that needs human judgement | A possible sanctions match stops the run and goes to compliance |
| Limit | A turn or spend cap is reached | After 12 tool turns, hand over what the agent found |
Autonomy has a price, which Anthropic's guidance names: higher costs, and errors that compound, as one wrong assumption early carries through every later step. It recommends extensive testing in sandboxed environments with appropriate guardrails. Anjali's agent gets read-only tools plus one that saves a draft. With no tool that can email a customer or book space, a wrong turn ends in a draft that a person reads, never in a sent quote.
1.3.5 Choosing the pattern: the requirement decides
Bruno's question is fair: if the agent can handle the hardest 10%, why not give it everything? Because a pattern is chosen by the requirement it must meet, not by the hardest case it could handle. Anthropic's advice is to find the simplest solution possible and add complexity only when needed, since agentic systems often trade latency and cost for better task performance. Five requirements decide it.
- Predictability. Can you write the steps down before the request arrives? If yes, a workflow; if each step depends on what the last one found, an agent.
- Path variation. How often does the path change from one request to the next? A path that a fifth of requests share is worth a branch of its own; one that differs every time is an agent's job.
- Error tolerance. What does one wrong answer cost? Low tolerance favours fixed steps, gates in code and a person before the action.
- Cost and latency. Every extra call adds both, and an agent adds them unpredictably, so a tight reply-time target favours the fewest calls that meet the bar.
- Oversight and audit. A workflow runs the same steps for every case; an agent needs a step-by-step log and checkpoints to match.
| Pattern | When it wins | What it costs |
|---|---|---|
| One augmented LLM call | One step with the right context does the job, the input is well defined, and the lowest latency and cost matter | No second look: quality rests entirely on the retrieval and the prompt |
| Workflow | The steps are known in advance, the path rarely changes, and audit needs the same steps every time | More calls and latency than one call; someone maintains the paths; inputs no path anticipated fall through |
| Agent | The steps cannot be predicted or hardcoded, and you can bound what the agent may touch | Higher, variable cost and latency; compounding errors; harder to test and audit; needs limits and checkpoints |
Then Anjali let the evidence decide. She ran all three designs over the 400 labelled requests as an eval: fixed inputs with known right answers, scored the same way for every design. On standard lanes, the single call and the agent were equally accurate, but the agent used about eight times the tokens and took 40 seconds a reply instead of five. On special cargo, the workflow clearly beat the single call and matched the agent at a fraction of its cost. Only on unusual requests did the agent earn its keep: the workflow could only pass them to a person, while the desk accepted over half of the agent's drafts as written.
So the answer is a hybrid, and hybrids are normal: Anthropic presents these patterns as common building blocks to shape and combine, not as prescriptions. The final design is a routing workflow whose third branch is a bounded agent, recorded with its evidence, the rejected option and when to revisit it.
Decision 07: architecture for automated quote replies.
Requirement: every sent quote matches the rate card or is approved by the pricing desk; standard-lane replies within 15 minutes; a record of which steps ran for every quote.
Evidence: 400 labelled past requests run through three designs (one augmented call, a routing workflow, a tool-using agent), compared on accuracy, time and tokens per request type.
Decision: a routing workflow. Standard lanes take one augmented call with the lane's rates retrieved and the price checked in code. Special cargo takes a prompt chain with a required-fields gate and desk approval. Unusual requests go to a bounded agent that drafts for the desk.
Rejected: one agent for every request. It matched the workflow on standard lanes at about eight times the tokens and 40 seconds a reply, and its varying steps make the audit record harder to read.
Costs accepted: an unusual request costs far more than a standard quote and still needs a person; the router is one more call to maintain and evaluate.
Revisit when: the unusual share passes 15%, or a new model changes the eval results.
1.3.6 Frameworks or direct API calls: what the abstraction hides
Once the pattern is chosen, a second decision follows: write it against the API yourself, or adopt something that runs the loop for you? The options run from plain Messages API calls, through the client SDKs' tool runner (in beta) and the Claude Agent SDK, to Claude Managed Agents (also in beta). There Anthropic runs the harness (the loop and tool execution around the model), in a sandbox it manages or one you host yourself. Third-party frameworks, such as the Strands Agents SDK by AWS, sit alongside them.
They all save work, and they all hide something. Frameworks handle the chores (calling the model, defining and parsing tools, chaining calls), but Anthropic warns that their extra layers can obscure the prompts and responses underneath and tempt you into needless complexity. A framework is like an automatic gearbox: easier to drive, until you need to know which gear you are in. Anthropic's advice is to start with the API directly, since many patterns take a few lines, and to understand the code beneath any framework you adopt.
Defaults are where hidden behaviour bites. The tool runner loops until Claude replies without a tool call, or until max_iterations if you set it; the Agent SDK has no turn or spend limit unless you set max_turns or max_budget_usd. Keelhaven's agent prototype ran on the tool runner and did well in a demo. Then two production requirements moved it to a manual loop: an audit entry for every tool call, and a hard stop the moment screening returns a possible match. The tool runner's documentation points to the manual loop for exactly such needs: human approval, custom logging and conditional execution.
Here is the loop Anjali wrote. Every API response carries a stop_reason, and the value tool_use means Claude is asking your code to run a tool. Look at the three lines marked STOP and the audit_log call: each is a decision that now sits in your code, where you can see and test it.
def run_agent(case_id, messages):
for turn in range(MAX_TURNS): # STOP 1: a turn limit you chose
response = client.messages.create(model=MODEL, max_tokens=4096, system=SYSTEM,
tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use": # STOP 2: finished, or any other stop
return to_pricing_desk(case_id, messages)
results = []
for call in (b for b in response.content if b.type == "tool_use"):
output = run_tool(call.name, call.input) # read-only tools, plus save_draft
audit_log(case_id, turn, call.name, call.input, output) # every step, in your own store
if call.name == "screen_parties" and output["status"] == "possible_match":
return to_compliance(case_id, messages) # STOP 3: a rule in code, not a prompt
results.append({"type": "tool_result", "tool_use_id": call.id,
"content": json.dumps(output)})
messages.append({"role": "user", "content": results}) # the environment's answer
return to_pricing_desk(case_id, messages) # limit reached: hand over, never guess
| Option | When it wins | What it costs |
|---|---|---|
| Direct API calls | Chains, routers and short loops; any step you must log, gate or explain; Anthropic's suggested starting point | You write the loop, the retries and the state handling yourself |
| SDK tool runner (beta) | A plain tool loop you want working quickly: it runs the tools, returns the results and keeps the conversation | The steps sit inside a helper; for approval, custom logging or conditional execution, the docs send you to a manual loop |
| Claude Agent SDK | Agents that need Claude Code's built-in tools, permissions, hooks, sessions or subagents, in a process you operate | A larger runtime with defaults to learn, such as no turn or spend limit until you set one |
| Claude Managed Agents (beta) | Long-running or asynchronous work where you would rather not build the loop, sandbox and tool execution yourself | Conversation history, sandbox state and outputs are stored on Anthropic's side until you delete the session, so it is not currently eligible for Zero Data Retention or HIPAA business associate agreement coverage |
| Third-party frameworks | A team already standardised on one, or a visual builder for simple chains | Extra layers that can hide prompts and responses, and an easy path to needless complexity |
Keelhaven's router and chains call the API directly: each is a few dozen lines, and every prompt and response lands in Keelhaven's own logs.
1.3.7 The exam traps
Nearly every trap here reaches for more autonomy or machinery than the requirement asks for, or trusts a prompt to do a control's job.
- ✗ Building an agent because the request said "agent". ✓ Start with the simplest pattern that meets the requirement, and add steps or autonomy only when evals on real inputs show the simpler one falls short.
- ✗ Using an agent for a fixed, well-defined sequence. ✓ Run it as a workflow. With known steps, autonomy only adds cost, latency and variation.
- ✗ Growing a routing workflow into an ever-larger decision tree for open-ended cases. ✓ Give the branch whose path cannot be predicted to a bounded agent or to a person, and keep the workflow for the rest.
- ✗ Trusting the prompt to stop the agent or to make it ask first. ✓ Set turn and spend limits, checkpoints and narrow tools in code. A prompt asks; a limit enforces.
- ✗ One pattern for the whole system. ✓ Choose per request type. A routing workflow with an agent on one branch is a normal, defensible design.
- ✗ Adopting a framework without knowing what it does underneath. ✓ Start with direct API calls and set any framework's limits yourself; its defaults become your production behaviour.
1.3.8 Put it together: compare three designs on one inbox
You now have the whole decision: the building block, the workflow patterns, the agent loop and its exits, the requirements that choose between them, and what a framework hides. The quickest way to make it stick is to run all three designs on one small inbox and let the numbers argue.
The rest of Domain 1 builds on this choice. Multi-agent systems (1.4) take orchestrator-workers further, when the workers become agents with loops of their own. Decomposition techniques (1.5) are how you cut a task into the steps a chain or an orchestrator runs. And when your agent must work with one that another team or company runs, agent-to-agent protocols (3.7) decide how they talk.
Key takeaways
- ✓ The augmented LLM, one model call with retrieval, tools and memory, is the building block of every pattern, and for many tasks it is the whole design.
- ✓ A workflow runs model calls along code paths fixed in advance; prompt chaining, routing, parallelisation (sectioning, voting), orchestrator-workers and evaluator-optimizer each fit one shape of task.
- ✓ An agent lets the model direct its own tool use in a loop, grounded by real tool results, until it finishes or a stop condition you set in code fires.
- ✓ Predictability, path variation, error tolerance, cost and latency, and oversight and audit decide the pattern; start with the simplest that meets them and add complexity only with evidence.
- ✓ Choose per request type: a routing workflow with a bounded agent on its one unpredictable branch is a normal, defensible design.
- ✓ Frameworks save plumbing but hide prompts, responses and defaults; start with direct API calls, know what any framework does underneath, and set its limits yourself.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
33 CCAR-P questions on Domain 1, free
Every question in the bank is tagged to a domain, so you can drill 33 questions on Solution Design & Architecture alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 1 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.