Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 1 › Lesson 1.4

CCAR-F · Domain 1 · 27% of the exam · Lesson 1.4 · 20 min read

Workflow enforcement and handoff: guarantees, not good intentions

Why a prompt saying "verify first" still fails, how a prerequisite gate blocks process_refund until get_customer succeeds, and what a handoff must contain.

Written against task statement 1.4 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

1.4.1 Why "always verify the customer first" is not enough

Picture a bank branch with a sign on every counter: "Check ID before paying out." Most tellers follow it most days. Then one Friday the queue is long, a regular says "you know me", and the teller pays out without looking at a card. The sign was clear. It was still only a sign. The branch that never has this problem is the one whose cash drawer will not open until an ID has been scanned.

Your support agent faces the same choice. Its system prompt (the standing instructions it reads in every conversation) says, in bold: "Always call get_customer to verify identity before looking up an order or issuing a refund." Yet once in a while it does not. A customer writes "Hi, this is Dana Morales, order 8812 arrived cracked, please refund it". The agent, reading a confident name and an order number, calls lookup_order with the name it was given and moves on to process_refund. If there are two Dana Moraleses, the wrong account gets the money.

The message made verification feel unnecessary, so the model skipped it. That is not a badly written prompt. It is what a prompt is: guidance the model weighs, not a wall. The locked drawer has a name: programmatic enforcement, a check in your own code that sits between the model's request and the tool. It refuses to run lookup_order or process_refund until get_customer has returned a verified customer id. This lesson covers when you need that guarantee and how to build it. It also covers two moves around it: splitting a message that raises three problems at once, and handing a case to a human who cannot see the conversation.

1.4.2 Prompt-based guidance versus programmatic enforcement

A rule about the order of steps can live in two places. It can live in the prompt, as an instruction the model reads. Or it can live in your code, as a check that runs whether the model remembers the instruction or not. Which place you choose decides how often the rule holds.

Prompt-based guidance is the instruction: "Verify identity before any financial operation." The model reads it on every turn and follows it nearly every time. Nearly is the word that matters. The model picks each tool call by reasoning over the whole conversation, and a persuasive message, a long history or an unusual phrasing can tip that reasoning the wrong way. So prompt instructions alone have a non-zero failure rate: small, but never zero.

Programmatic enforcement is the check in code. Before your code runs the tool the model asked for, it asks: has get_customer returned a verified id yet? If not, process_refund does not run, and the model gets an error explaining why. The model still chooses WHAT to ask for; your code decides WHETHER it runs. It comes in two forms. A prerequisite gate is the "has step A completed?" check we build in this lesson. A hook is code that the Claude Agent SDK (Anthropic's toolkit for building agents) runs at fixed moments in the loop, such as just before a tool call.

Where the rule lives decides how often it holds

Prompt says "verify first" probabilistic

Model reads the instruction
Model weighs it against the message
Usually calls get_customer first
Sometimes skips itnon-zero failure rate

Gate blocks the call

Model asks for process_refund
Code checks: verified id present?
No: call refused, error returned
Yes: call runs
An instruction in the prompt is one input to the model's reasoning; a gate in your code runs before every tool call and cannot be reasoned around.

1.4.3 When deterministic compliance is required

If gates are so reliable, why not gate everything? Because a gate is a rule written in advance, and the point of an agent is to handle what nobody wrote in advance. Every gate removes a little of the model's judgment. The real skill, and the one the exam tests, is knowing WHICH rules need a guarantee.

The test is the cost of one failure. Ask: if the model skipped this step once in a thousand conversations, what would happen? If the answer is "a slightly worse reply", the prompt is the right tool and a gate is over-engineering. If the answer is "money leaves the company, private data reaches the wrong person, or an account is changed for good", one in a thousand is one too many. That rule needs deterministic compliance: the same situation produces the same correct behaviour every time, with no probability involved.

The rule Cost of skipping it once Where it belongs
Greet the customer by name A colder reply Prompt
Apologise before explaining A slightly worse tone Prompt
Look up the order before quoting a delivery date One wrong estimate, easily corrected Prompt (a gate is optional)
Verify identity before lookup_order Someone else's order details disclosed Code
Verify identity before process_refund Money sent to an unverified person Code
Refunds above a limit go to a human An unauthorised large payment Code

Memorise the test behind the table, not the rows. Identity verification before a financial operation is the textbook case. Anything with the same shape (financial, legal, privacy-related or irreversible) belongs in code; style, tone and phrasing belong in the prompt. For Dana's refund that settles it. "No lookup_order or process_refund until get_customer has returned a verified customer id" protects both her order details and the company's money, so it goes in code.

1.4.4 Building the prerequisite gate

Now let's build the gate. It lives inside the code that runs tool calls: before running a gated tool, it confirms that the prerequisite step has completed, and if not, it refuses the call and tells the model why. It needs three things: a record of what has completed, a list of gated tools, and a way to refuse.

The record is a small piece of state your code keeps next to the conversation, here verified_customer_id, empty at the start of every conversation. The gated tools are lookup_order and process_refund. The refusal is an ordinary tool result with is_error set to true and a message that says what to do instead. That is the documented way to tell the model a tool call did not succeed. The model reads the message and, in practice, calls get_customer next. Even if it did not, nothing bad could happen: the refund stays blocked.

The gated refund path

Customer message"refund order 8812"
Model asks process_refundskipping verification
Gate checks stateverified id? no
Error result returned"verify the customer first"
Model calls get_customergate records the id
Model asks process_refund againid present, the call runs
The gate sits between the model's request and the real tool. Until get_customer has returned a verified id, requests for lookup_order and process_refund are refused with an error the model can act on.

Here is the gate inside the function that runs tools in a hand-written loop. Look at the two commented lines: the check that refuses the call, and the line that unlocks the gate only when get_customer reports a verified customer.

state = {"verified_customer_id": None}        # reset for every conversation
GATED = {"lookup_order", "process_refund"}

def run_tool(name, args):
    if name in GATED and state["verified_customer_id"] is None:
        return {"is_error": True,                          # the GATE refuses
                "content": "Blocked: call get_customer and verify identity first."}
    result = backend[name](**args)                          # the real backend call
    if name == "get_customer" and result["verified"]:
        state["verified_customer_id"] = result["customer_id"]   # UNLOCK the gate
    return {"content": result}
# your loop wraps whatever run_tool returns in a tool_result block and sends it back

Notice how little the gate does. It does not script the order of the other tools or stop the model from asking the customer a question. It is the smallest possible fence around the one step that must never be skipped. And it unlocks on the RESULT of get_customer, not on the fact that it was called: a lookup that failed or came back unverified leaves the gate shut. If you build on the Agent SDK, which runs the loop for you, the same check goes in a hook that runs just before each tool call. Same idea, different place.

1.4.5 Three problems in one message

Real support messages rarely raise just one issue. Dana's full message reads: "Order 8812 arrived cracked, you charged me twice for it, and now the app says my account is locked." That is three problems with three different causes. The failure to avoid is an agent that fixes the cracked item, replies, and drops the other two, or blends all three into one muddled answer.

The pattern that avoids it has three moves. First, decompose the message into distinct items, each with its own question: is the item damaged and refundable, was Dana charged twice, why is the account locked? Second, investigate the items in parallel with shared context. They do not depend on each other, so there is no reason to finish one before starting the next. But they all need the same facts: the verified customer id and the order number. Third, synthesise the findings into one unified resolution, a single reply that answers all three.

One message, three investigations, one reply

Cracked item

Shared: customer id, order 8812
lookup_orderdelivery record
Damaged on arrival: refund $62

Charged twice

Shared: customer id, order 8812
lookup_orderpayment records
$184 charged twice: over the limit

Account locked

Shared: customer id
get_customeraccount status
Locked after failed sign-ins
Every investigation starts from the same shared facts and returns its own finding; the findings are then combined into one resolution.

"In parallel" can happen in two ways. In a single agent, the model can request several tool calls in one turn; your code runs them and returns all the results together in the next message. In a design with a coordinator (an agent that hands sub-tasks to other agents), the coordinator starts one subagent per item, and the Agent SDK can run them at the same time.

Either way, the shared context has to be passed on purpose. A subagent starts with a fresh conversation and does not see the coordinator's history. The only case facts it gets are the ones the coordinator writes into its instructions, so leave out the verified id and each investigation starts by asking who the customer is.

Dana then gets one reply. The cracked item is refunded, the duplicate charge has gone to a colleague because $184 is above the agent's refund limit, and the account lock clears when she resets her password. Three answers, one message, nothing dropped, and the one item that needs a human does not hold up the other two.

1.4.6 The handoff a human can act on

Sooner or later the agent reaches the edge of what it may do: a refund above its limit, a policy exception, a customer who asks for a person. It calls escalate_to_human, and the question that matters is WHAT it puts in that call. Here is the constraint that shapes the answer: the human who picks up the case cannot see the conversation transcript. They see what the handoff contains, and nothing else.

Think about what that means in the middle of Dana's case. The agent has verified her, refunded the cracked item, and found that the payment provider retried a charge that had timed out, so the order was charged twice. If the escalation says "customer has a billing issue, please help", the human starts from zero: re-verifies Dana, re-reads the order, rediscovers the double charge. Worse, they might refund the wrong amount, or refund the cracked item a second time.

A structured handoff summary prevents that. It is a fixed set of fields the agent fills in every time it escalates, so the human can act at once without the transcript. Using the same fields on every case is what makes it a protocol: the human team always knows where to look. Four fields are essential.

Field What goes in it Why the human needs it
Customer The verified id from get_customer (not the typed name), plus the details on file They can trust identity without re-verifying
Root cause What actually went wrong, established from tool results They act on the diagnosis instead of redoing it
Refund amount The exact figure and what it covers They approve or adjust a number, not a vague request
Recommended action What the agent proposes, and what it has already done They decide quickly and do not repeat a completed step

Root cause deserves emphasis. "Customer says charged twice" is a symptom; the root cause in the example below is a diagnosis, established from the payment records. To make the protocol hard to skip, make the four fields required inputs of escalate_to_human. A call that leaves one out is then invalid, and your code can refuse it with the same kind of error result as the gate.

1.4.7 The exam traps

Every wrong answer on this task statement either trusts the model to do what code should guarantee, or hands a human less than they need.

  • ✗ Strengthening the prompt when a mandatory step is being skipped. ✓ Add a prerequisite gate (or hook) that blocks the downstream call until the step has completed. Capitals, repetition and few-shot examples (sample conversations that show the right order) lower the failure rate; a financial step needs it at zero.
  • ✗ Gating every rule, including tone and phrasing. ✓ Gate what is financial, legal, private or irreversible; leave style to the prompt. Over-gating turns the agent back into a script that fails on requests nobody anticipated.
  • ✗ Unlocking the gate because get_customer was called. ✓ Unlock on the RESULT, a verified customer id. A lookup that failed or returned "unverified" must not open the refund.
  • ✗ Answering the first issue in a three-issue message and stopping. ✓ Decompose into items, investigate in parallel with shared context, and reply once with all three addressed.
  • ✗ Sending subagents off with no shared context. ✓ Pass the verified customer id and case facts into each investigation, or each one starts by re-identifying the customer.
  • ✗ Escalating with "customer needs help" or a pasted transcript. ✓ A structured summary: customer id, root cause, refund amount, recommended action. The human cannot see the transcript, and a wall of text is not a summary.

1.4.8 Put it together: gate the refund and watch the gate hold

You now have every piece: guidance versus enforcement, the cost test for which rules need a guarantee, the gate itself, the three-issue split and the handoff. The gate is the part to feel for yourself: build it, remove it, and watch the failure it was protecting you from.

The gate you built by hand is what hooks (1.5) give you inside the Agent SDK: a fixed place, just before or just after each tool call, where a check like this one lives. The three-issue split is one instance of task decomposition (1.6): which parts are independent, which must be sequential, and what context each part needs. And a handoff carries state across a boundary between agent and person; sessions (1.7) handle the case where the boundary is time.

Key takeaways

  • ✓ A prompt instruction about workflow order is guidance the model weighs; it lowers the failure rate but leaves it above zero.
  • ✓ Programmatic enforcement (a prerequisite gate or a hook) runs in your code before the tool executes, so the model can request a step but cannot bypass the check.
  • ✓ Deterministic compliance is required when one skipped step costs money, exposes private data, breaks a law or cannot be undone; identity verification before a financial operation is the textbook case.
  • ✓ A prerequisite gate blocks lookup_order and process_refund until get_customer has RETURNED a verified customer id, and refuses with an instructive is_error result the model can act on.
  • ✓ A message raising several concerns is decomposed into items, investigated in parallel with shared context, and answered in one unified reply.
  • ✓ A handoff to a human who cannot see the transcript is a structured summary: customer id, root cause, refund amount, recommended action.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

96 CCAR-F questions on Domain 1, free

Every question in the bank is tagged to a domain, so you can drill 96 questions on Agentic Architecture & Orchestration alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 1 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources