Home › Study guides › CCAR-P › Domain 5 › Lesson 5.3
CCAR-P · Domain 5 · 14% of the exam · Lesson 5.3 · 22 min read
Human-in-the-loop validation: putting review where the risk is
Where a person should approve, spot-check or monitor Claude's output, how to make that review real, and how to build and record approval gates in agents.
Written against objective 5.3 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
5.3.1 Why "an officer reviews every memo" failed
Ashenford Bank's business-lending team writes about 150 credit memos a week. A credit memo is the document a credit officer decides a loan on: the borrower, its financial statements and account history, its ability to repay, and the security and covenants that protect the bank. A covenant is a promise written into the loan, such as keeping operating cash flow at least 1.25 times the year's debt payments. In a pilot, Claude drafted memos in minutes instead of an analyst's day.
Obafemi, the chief credit officer, set one rule for the pilot: an officer reads and approves every memo in full. Three months in, Ingeborg, the solution architect, pulled the numbers. The queue was three days deep, the median review of a 25-page memo took seven minutes, and 99.6% of memos were approved. Then an internal audit found a memo that had carried last year's cash-flow figure into this year's covenant test, hiding a breach. An officer had approved it in four minutes.
That officer was not unusually careless. When a system is almost always right, busy people checking it start accepting its output without looking, and one Approve button under 25 pages invites exactly that. The pilot had a person in the loop on paper and a rubber stamp in practice, and the memos that deserved an hour got the same seven minutes as a routine renewal.
Human-in-the-loop validation is the design of where people check, approve or oversee what an AI system produces. For each output it settles whether a person sees it, when (before it takes effect or after), with what evidence, and what gets recorded. Ingeborg's redesign keeps officers in charge of every lending decision, but spends their attention where an error would be costly, hard to undo, or not the bank's to automate.
Review everything, or review by risk
Pilot: officers approve every memo
Redesign: review by risk
5.3.2 Five ways to put a person in the loop
Here is the belief that trips people up: human-in-the-loop means a person approves each output. That is one pattern of five, and often not the best one. The patterns differ in WHEN the person acts, before or after the output takes effect, and in how many cases they see.
| Pattern | When it wins | What it costs |
|---|---|---|
| Approval before action | The output is hard to undo, affects a person, or policy names who must decide | Latency on every case and reviewer hours; rubber-stamping when volume is high |
| Review after the fact | The action is cheap to reverse and a short-lived error does little harm | Errors reach the world first; needs a fast correction path |
| Sampled review | You need to measure quality across many low-risk outputs | Protects no single case; the sample must be random or stratified to mean anything |
| Exception-based escalation | Most cases are routine, and risk or doubt shows in signals code can compute | Only as good as its triggers; misses what no trigger describes |
| Monitoring the running system | High volume; a person watches rates and trends and can pause or reroute | Sees patterns, not single cases; needs alert thresholds and a working stop switch |
Memorise the five names and the middle column; most distractors apply one of these patterns to the wrong risk.
Ashenford uses all five. The risk rating, recommendation, repayment analysis and covenant terms need an officer's approval before a memo reaches the lending system. A covenant breach or a large exposure escalates to a senior decision-maker. Quality assurance reads one memo in ten in full. The summary the agent posts to the relationship manager's customer notes is reviewed after the fact, because a wrong note is fixed in a minute. And Obafemi watches a daily dashboard of edit and escalation rates.
Don't underrate monitoring. Anthropic's study of millions of agent interactions found that experienced Claude Code users auto-approve more often than new users, yet also interrupt more often: they shift from approving each action to watching and stepping in. Its conclusion is that effective oversight doesn't require approving every action, but being in a position to intervene when it matters. A lifeguard approves nobody's swimming stroke, yet someone is watching and can act.
5.3.3 Choosing the tier: reversibility, people and policy
"Which outputs need a person?" has no answer until you ask three questions of each output. Can it be undone, and at what cost? Does it affect a person's money, rights or opportunities? Does law, regulation, internal policy or a provider's usage policy require a qualified person to decide? A yes pushes the output towards approval before action; the size of the stake turns the dial further.
For Ashenford the third question settles the top tier: credit policy requires a named officer's decision on every loan. Anthropic's Usage Policy adds a floor of its own. It counts loan approvals and creditworthiness among its high-risk use cases, and where such decisions directly affect individuals or consumers, it requires a qualified professional to review them before they are finalised. The architect's job is to turn each obligation, internal or external, into a gate.
One memo, three tiers of review
Officer decides each line
Spot-checked
1 memo in 10 read in full by quality assurance
Escalated automatically
to a senior officer or the credit committee
Escalation triggers come in two kinds. Risk signals say the stakes are high: a covenant breach, a borrower on the bank's watch list, or group exposure above the officer's delegated authority. Group exposure is everything lent to the borrower and its related companies; delegated authority is the most an officer may approve alone. Low-confidence signals say the draft may be wrong: a figure that doesn't reconcile with the statements, a claim with no source, two independent drafts that disagree, or the model saying it lacks the information. Anthropic's hallucination guidance backs the last two: treat inconsistency across repeated runs as a warning sign, and give Claude permission to say it doesn't know.
Missing from both lists: the model rating its own confidence. That score is more generated text, checked against nothing, and a draft built on last year's cash-flow figure can arrive with a high one. Low confidence is something your code measures, not something the draft reports about itself. The model's "I don't know" may add a case to the queue; its self-rated confidence must never remove one. Ashenford's code recalculates the covenant test from the statements and reads exposure from the lending system, never from the draft. Descriptive sections that pass these checks get sampling alone, which frees officers for the lines that decide the loan.
5.3.4 Making the review real
Tiering alone would fail if officers still faced 25 pages and one button. The enemy is automation bias: the tendency to accept a machine's suggestion without checking it, strongest when the machine is usually right and the reviewer is busy. You counter it by design, making checking the easy path and blind approval the awkward one. A pre-flight checklist works this way: each item demands an action on one specific thing, not a signature under "aircraft checked".
Four design moves do most of the work. Show the evidence: each figure sits beside the passage it came from, which the Citations feature supports by returning the exact cited text and, for PDFs, page numbers. A citation proves the passage exists, not that it supports the sentence. Show what changed since the last approved memo, for renewals. Surface the doubt: flags sit at the top, where nobody can scroll past them. Demand a decision per risky line: accept, edit or reject, with a reason for anything but accept. Look at line 5.2 on Ashenford's review screen, which has no source and cannot be accepted as it stands.
Memo LN-24-0917, version 3. Section 5: Covenants. Decide each line.
5.1 "Operating cash flow covers debt payments 1.18 times, below the 1.25 minimum in the loan agreement." Source: FY2025 audited statements, page 14, cash flow statement. Recalculated by the covenant check: 1.18. Flag: ESCALATED (covenant breach), routed to a senior officer.
5.2 "No other covenant has been breached in the last 12 months." Source: none found. Flag: UNSUPPORTED. Confirm against the covenant register or strike the line.
Decision per line: Accept / Edit / Reject. A reason is required for Edit and Reject.
Then measure whether the review works. Ashenford seeds a few memos each month with known errors, such as a wrong covenant level, and tracks how many officers catch, along with time per review. A 99% approval rate in under a minute is a symptom, not a success.
Workload is part of the design, and you size it like any queue. At 150 memos a week, with about six decided lines and ten minutes per memo, officers need some 25 hours a week, plus escalations, to meet a one-business-day turnaround. When the queue exceeds capacity, memos wait and Obafemi is alerted. The overflow path is never automatic approval, and never a thinner screen.
5.3.5 Building the gate into the agent
A gate written in the prompt is not a gate: "always wait for approval" is guidance the model usually follows, and a control must hold every time. Claude decides WHAT it wants to do; your code decides WHETHER it happens. Ingeborg's agent runs on the Claude Agent SDK, which gives her two places to enforce a gate; the third lives in the lending application. The choice turns on who approves and how long they take.
| Mechanism | When it wins | What it costs |
|---|---|---|
Permission prompt: the SDK's canUseTool callback shows the call to a person, who allows or denies it |
The approver is in the session, such as the analyst running the agent, and answers within minutes | The run stays paused while the callback waits; calls already approved by a rule or permission mode never reach it |
Hook: a PreToolUse hook returns allow, deny, ask or defer for each call |
The rule depends on the call's inputs and must run on every call, whatever the rules and mode say | Code you own and test; defer, which ends the run with the call saved for a later resume, works only when Claude makes one tool call in that turn |
| Pending-approval state: the agent's tool only places the output in a review queue | The approver is someone else, takes hours, or the approval must outlive the process and be audited | You build the queue, the review screen and the state transitions; the agent's job ends at "submitted" |
At Ashenford, the analyst running the agent confirms an email asking a borrower for missing documents in seconds, so a prompt fits. The officer's decision takes hours and must be audited, so it is a pending-approval state in the lending application. A hook enforces what must hold on every call. Look at the submit_memo branch: it reads the application's own memo_store instead of trusting Claude's account of the memo, and even its allow only parks the memo for an officer.
from claude_agent_sdk import ClaudeAgentOptions, HookMatcher
def decide(decision: str, reason: str) -> dict:
return {"hookSpecificOutput": {"hookEventName": "PreToolUse",
"permissionDecision": decision, "permissionDecisionReason": reason}}
async def lending_gate(input_data, tool_use_id, context):
tool, args = input_data["tool_name"], input_data["tool_input"]
if tool == "mcp__lending__email_borrower": # leaves the bank: a person confirms
return decide("ask", "External email: check the recipient and wording")
if tool == "mcp__lending__submit_memo":
memo = memo_store.get(args["memo_id"]) # the application's record, not Claude's word
if memo.missing_sections or memo.uncited_figures:
return decide("deny", "Incomplete: missing sections or figures with no source")
return decide("allow", "Complete") # the tool only queues it as PENDING_APPROVAL
return {} # every other call: the normal permission flow
options = ClaudeAgentOptions(
can_use_tool=ask_analyst, # the analyst's screen answers each "ask"
hooks={"PreToolUse": [HookMatcher(hooks=[lending_gate])]},
)
Two documented details matter here. A call already approved by a rule or permission mode never reaches canUseTool, so a check placed only there can be silently skipped, while a PreToolUse hook runs first and sees every call. And a deny reason goes back to Claude, which can tell the analyst what is missing, while an ask reason is shown only to the person. Note also that the agent has no tool that files a memo or records a credit decision, so only the lending application acts, after an officer approves.
5.3.6 Recording the decision: audit trail and feedback
It is tempting to treat the approved memo as the record. But an auditor will ask who approved the loan, which version they saw, which flags were on screen and what they changed. The approved memo answers none of that, and it discards what officers corrected.
Ashenford writes one decision record per memo version to an append-only store, kept as long as the bank's record-keeping policy requires. Here is the record for the memo on the review screen above. Look at the first line, which pins a fingerprint (hash) of the exact text reviewed, and at 5.2, which keeps the officer's reason and new wording.
Record DR-0917-3. Memo LN-24-0917, version 3. SHA-256 of the reviewed text: 9f2c...e41a. Prompt v12 on the pinned model named in release R-2026-09.
Shown to the reviewer: 31 citations; flags on 5.1 (ESCALATED, covenant breach, recalculated at 1.18) and 5.2 (UNSUPPORTED).
5.1 Accept. Decided by senior officer S-118 after escalation.
5.2 Edit. Reason: the covenant register shows financial statements delivered 21 days late in March 2026. New text: "One other breach in the last 12 months: late delivery of financial statements, March 2026."
Recommendation: Edit, "renew" changed to "renew subject to a covenant waiver and quarterly monitoring". Reason: two breaches in 12 months. Decided by credit officer C-204, countersigned by S-118.
Time on screen: 14 minutes. Written on 2026-09-14 at 14:05.
The same record feeds improvement. Edited and rejected lines become cases in the evaluation set, so the next prompt change is tested against real mistakes. A high edit rate on one section tells Ingeborg that the prompt or the data feed needs work, not more reviewers. Escalations that officers routinely wave through show which triggers are too sensitive.
Evidence moves sections between tiers in both directions. If quality assurance finds the industry overview sound in almost every sample for a quarter, the sample rate can drop. If the edit rate on covenant lines jumps after a model upgrade, Obafemi can send every memo back to full review until the cause is found. Tiers are a hypothesis you keep testing.
5.3.7 The exam traps
Every trap here either puts review in the wrong place or trusts something nobody checked.
- ✗ Sending every output to a person "to be safe". ✓ Tier by reversibility, impact on people and what policy requires. Blanket review builds a queue and breeds rubber-stamping.
- ✗ Sampling after the fact for an action that can't be undone. ✓ Approve before action. A sample measures quality across many outputs; it cannot stop the one wrong payment, notice or decision.
- ✗ Letting the model's self-rated confidence decide what skips review. ✓ Escalate on risk and low-confidence signals your code computes: thresholds from systems of record, figures that don't reconcile, claims without sources, disagreement between runs.
- ✗ An Approve button under a summary, or an "I have reviewed this" box. ✓ Put sources, changes and flags on screen and require a decision per risky line with a reason. An acknowledgement is not an inspection.
- ✗ Writing the approval rule into the system prompt. ✓ Enforce it in a hook, a permission callback or an application state the model cannot skip.
- ✗ Adding an approval step to a capability the role never needs. ✓ Remove the capability. Approval is a compensating control for actions the role must keep; removal is preventive, and nobody can approve a call that can't happen.
5.3.8 Put it together: design the review for one workflow
You now have the whole method: patterns, tiers, computed triggers, real review, gates in code and a decision record. The quickest way to make it stick is to build a small gate, feed it the wrong signal, watch it miss a breach and restore it.
Compliance (5.4) settles which decisions the law reserves for people and how long decision records must be kept. Ethical AI (5.5) checks outcomes by group, for drafts and reviewers' decisions alike, and adds appeal routes. And the turnaround you promise the business becomes an SLA managed with stakeholders (6.3).
Key takeaways
- ✓ A person in the loop is a control only when the review is placed by risk, shows the evidence and is recorded; reviewing everything tends to produce a queue and a rubber stamp.
- ✓ The five patterns are approval before action, review after the fact, sampled review, exception-based escalation and monitoring, and they differ in when a person acts and how many cases they see.
- ✓ Tier each output by reversibility, impact on people and what law or policy requires, and escalate on risk and low-confidence signals your code computes, never letting the model's self-rated confidence clear a case.
- ✓ Make review real with sources beside each claim, flagged doubts, a decision per risky line with a reason, seeded test errors, and a queue that waits rather than approving itself.
- ✓ Enforce gates in code: a permission callback for an approver in the session, a
PreToolUsehook for rules that must run on every call, and a pending-approval state for approvals that take hours. - ✓ Record every decision with the version and evidence it covered; the record is both the audit trail and the feedback that retunes prompts, evals, triggers and tiers.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
27 CCAR-P questions on Domain 5, free
Every question in the bank is tagged to a domain, so you can drill 27 questions on Governance, Safety & Risk Management alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 5 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.