Home › Study guides › CCAR-P › Domain 5 › Lesson 5.5
CCAR-P · Domain 5 · 14% of the exam · Lesson 5.5 · 23 min read
Bias, fairness and transparency when Claude helps decide about people
Where bias enters a Claude solution, how to test for it with counterfactual pairs and group breakdowns, which mitigations hold, and what people must be told.
Written against objective 5.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
5.5.1 Why a consistent shortlist can still be unfair
Kestrelmoor Staffing fills warehouse and office jobs at about forty client sites: night-shift pickers, forklift drivers, receptionists and payroll clerks. Around 3,000 applications arrive each week, and eleven recruiters cannot read them all before the best candidates accept other offers. Nnamdi, the operations director, wants Claude to read each CV and the answers to the application questions, then hand recruiters a shortlist for every role. Pernille, the solution architect on the engagement, is confident that Claude can do the reading. What worries her is something else.
Nobody at Kestrelmoor means to be unfair. The risk is that the system will be unfair anyway, quietly and at scale. A recruiter who marks down a CV for a two-year gap does it to a handful of people a week. A prompt that says "prefer continuous employment" does it to every parent, carer and person who has been ill, thousands of times, and every one of those decisions looks consistent. Consistency feels like fairness, but a system can apply the wrong rule perfectly.
Let's fix three terms first. Bias here means a systematic difference in how a system treats people that comes from something other than what the decision is about. Fairness is the property you design and test for: comparable people get comparable outcomes, and the criteria themselves are about the job. Transparency is what people are told and what is written down: that AI is involved, what it decides, where its limits are and how to challenge a result. None of the three is a setting on the model, so none of them can be handed to it.
5.5.2 Where bias gets in
Here is the belief that trips people up: bias is something the model has, so the fix is a better model. Models can carry bias. Claude learned language from human writing, which is full of associations about which jobs suit which people and what a "professional" CV sounds like. But in a system like Kestrelmoor's, the model is only one of six places where bias gets in, and the architect's own pipeline supplies the other five.
Think of a cake that comes out wrong: it might be the oven, but more often it is the recipe or the ingredients. Pernille walked Nnamdi's team through the design and found an example of each route.
| Where it enters | How it gets in | What it looked like at Kestrelmoor |
|---|---|---|
| The model | Tendencies learned from human writing, such as reading polished prose as competence | Higher scores for fluent, idiomatic English on warehouse roles that need only basic reading |
| The prompt and criteria | A criterion unrelated to the job, or a vague one the model fills with its own assumptions | "Prefer continuous employment" and "a young, energetic team", copied from a client brief |
| Few-shot examples | Examples teach the pattern of who was chosen, not only the output format | Three "strong candidate" examples for forklift roles, all men in their twenties |
| The retrieval corpus | Retrieved text brings its authors' assumptions into the context | Old job adverts retrieved as role descriptions, one asking for "strong lads" |
| Historical data | Past decisions used as labels or targets repeat past preferences | A plan to "shortlist people like the placements clients kept" |
| The evaluation set | Labels copied from past decisions score the old bias as accuracy, and missing groups hide failures | Last year's recruiter picks as the answer key, with almost no applicants over 55 |
Memorise the six routes; in a scenario, the symptom usually points to one of them. The last row deserves extra attention, because the evaluation set decides whether you ever see the other five. If your answer key is last year's shortlists, a system that copies last year's preferences scores well, and the test reports success.
5.5.3 Testing for fairness: change one thing, then count by group
How would Pernille know whether the shortlist is fair? Asking Claude is not a test: a model's account of its own reasoning is not a measurement of its behaviour. Anthropic's evaluation guide makes the point that even hazy qualities such as ethics and safety can be turned into measurable criteria. For fairness, three tests do the measuring, and each catches something the others miss.
The first is the counterfactual test. Take a CV, make a copy that differs in exactly one attribute, run both through the same prompt, and compare the results. It works like a wine tasting with swapped labels: pour the same wine into two glasses, change only the label, and if the ratings differ, the label did it.
The attribute can be a protected one, such as a name that signals gender or ethnicity. It can also be a proxy, a detail that stands in for a protected one, such as a graduation year for age or a career break for caring. Claude's outputs vary a little from run to run, so the test needs a no-change control, the identical CV run several times. A gap counts only when it is bigger than the control's spread.
This is the method Anthropic used in its 2023 research on discrimination in model decisions. The researchers generated 70 decision scenarios, such as whether to offer a job or pay out an insurance claim, with slots for age, race and gender, and varied only those. With no mitigation, Claude 2.0 favoured some groups and disfavoured others in some settings. Plain-language instructions reduced both, and two nearly eliminated them in those scenarios: saying that discrimination is illegal, and asking the model to answer as if no demographic information had been given. Borrow the method, not the result. The study measured one model on its own scenarios; your criteria, your applicants and today's model are untested until you test them.
In Pernille's harness, MODEL is the model Kestrelmoor plans to ship and RUBRIC is the scoring rubric. The call uses structured outputs: given a Pydantic class, messages.parse returns every reply as a valid score with its evidence, so no run is lost to a parsing error. Look at the control, which changes nothing, and the variant, which changes only the name.
from pydantic import BaseModel
import anthropic
client = anthropic.Anthropic()
class Assessment(BaseModel):
evidence: list[str] # one quote per rubric criterion
score: int # 0 to 12, from the rubric
def score(cv: str) -> int:
reply = client.messages.parse(model=MODEL, max_tokens=4000, system=RUBRIC,
messages=[{"role": "user", "content": cv}], output_format=Assessment)
return reply.parsed_output.score
for cv in BASE_CVS: # fictional CVs with a [NAME] slot
control = [score(cv.replace("[NAME]", NAME_A)) for _ in range(3)] # nothing changes: the noise
variant = [score(cv.replace("[NAME]", NAME_B)) for _ in range(3)] # ONLY the name changes
print(max(control) - min(control), sum(variant) / 3 - sum(control) / 3) # spread, then gap
One CV, three versions
Control identical CV, three runs
spread of 3: the noise
Name swapped
gap of 0: inside the noise
Caring break added
gap of 4: find the cause
The second test runs on real traffic: results broken down by group. Pernille compares the shortlist rate, the average score and the rate of recruiter overrides for each group, using the answers applicants give on Kestrelmoor's voluntary equal-opportunities form. That data is held apart from the screening input: the model never sees it, and only the analysis job does. An aggregate hides groups, so 91% agreement with recruiters overall can be 95% for one group and 70% for another. Agree the gap that triggers an investigation with HR and legal before launch; how the group data may be collected and kept is a compliance question.
The third test is a review of the criteria themselves. The first two check that the system applies its criteria consistently; neither can say whether a criterion belongs there at all. "Prefer continuous employment" passes a name swap perfectly and still excludes carers for a reason unrelated to the job. For each criterion, ask which job requirement it measures and whether someone who fails it could still do the work. Pernille's review struck out the continuity rule and limited English on warehouse roles to reading pick lists and safety signs. She put each struck rule to the client as a choice: name the job requirement it measures, or drop it.
5.5.4 Mitigations, from the criteria outwards
Once a test finds a gap, the tempting fix is one line in the prompt: "Do not discriminate." Keep the line, since Anthropic's research found that instructions like it can cut measured discrimination sharply. But an instruction is guidance you then have to verify, and it cannot repair a criterion that is unfair by design. The mitigations that hold are layered in the order a CV meets them: fix what goes in, constrain how the judgement is made, and control what happens after.
The controls in the order a CV meets them
Pernille's pipeline masks in code, before the call. Telling Claude to "ignore the name" leaves the name in the context; removing it means there is nothing to ignore. Where a client has a real need behind an attribute, replace the attribute with the fact the job needs. A client that asks for postcodes because night shifts start after the last bus gets a field computed in code, "can reach site 14 by 22:00: yes or no", instead of an address. Masking does not remove proxies, which is why the tests stay.
The rubric turns a vague judgement into a set of checks, each backed by evidence. Look at the last two lines of Pernille's rubric for one role: the do-not-use list, and the rule that Claude never outputs a rejection.
Role: night-shift warehouse operative, client site 14. Mark each criterion Met, Not met or Not stated, and quote the words in the application that support it. Each Met scores 3 points.
1. Shift pattern: available 22:00 to 06:00, four nights a week.
2. Manual handling: experience of, or willingness to, lift loads up to 20 kg.
3. Counterbalance forklift certificate: held, or booked before the start date.
4. Reading: can read pick lists and safety signs in English; do not reward style, spelling or fluency beyond this.
Do not use or infer: name, age, gender, ethnicity, nationality, religion, disability, family status, gaps in employment, address, photo, school names or hobbies. If something is not stated, record Not stated; never guess.
Output the four results with their quotes, then Shortlist or Refer to recruiter. Never reject: a recruiter decides every application.
| Mitigation | When it wins | What it costs |
|---|---|---|
| Job-related criteria, agreed per role | Always; every other control depends on them | Client conversations, and some client requests refused |
| Masking in code before the model call | The attribute is irrelevant and can be removed reliably | Proxies survive; recruiters need a way back to the person |
| Evidence-based rubric | Outcomes must be explained, audited and compared | Design time, longer outputs, less room for judgement calls |
| Anti-bias instruction in the prompt | A cheap extra layer on any decision prompt | Proves nothing alone; test it like any other change |
| Human review before the decision is final | Consequential decisions about people; required for high-risk uses | Recruiter time; a reviewer who sees only the shortlist becomes a rubber stamp |
| Outcome monitoring by group | Every live system that decides about people | Group data held apart, a lawful basis for it, and an owner who acts |
Memorise the first column; the costs are what you argue about when someone wants to drop a layer. The review row hides the subtlest trap. If recruiters see only Claude's shortlist, they can turn down a weak candidate Claude liked, but they never notice the strong candidate Claude left out, and wrongful exclusion is where unfairness usually hides. So every application reaches a recruiter with the quotes that make review quick, and each assessment is stored with the rubric version and model id so any outcome can be traced later.
5.5.5 Transparency, and what Anthropic requires for high-risk uses
A candidate Kestrelmoor turns down has three fair questions: was a machine involved, what did it decide, and how can I challenge it? Transparency is the design that answers them, and it has four parts. Disclose that AI is involved, in plain words. Scope it: say what the AI does and what it does not decide. Document the system and its limits. And provide an appeal route to a person who can change the result.
Anthropic's Usage Policy (the version in effect since 15 September 2025) turns part of this into requirements. It defines High-Risk Use Cases as consumer-facing uses in domains vital to public welfare and social equity. There are seven: legal; healthcare; insurance; finance; employment and housing; academic testing, accreditation and admissions; and media content generated and published automatically. The employment category names resume screening and hiring tools, so Pernille treats Kestrelmoor's shortlist as a high-risk use.
| Requirement | What the Usage Policy says today | Kestrelmoor's design |
|---|---|---|
| Human in the loop (high-risk uses) | For advice, recommendations or subjective decisions that directly affect people, a qualified professional in the field reviews the content or decision before it is disseminated or finalised | A recruiter confirms every shortlist and every referral |
| Disclosure (high-risk uses) | When outputs are presented directly to people, tell them AI helped produce the advice, decision or recommendation, at least at the start of each session | The notice on the application form and in every outcome message |
| Chatbots (any use) | A consumer-facing chatbot or external-facing agent tells users it is AI, not a human, at least at the start of each chat session | The candidate help chat opens by saying it is an AI assistant |
| No discrimination (all users) | Promoting discriminatory practices against people on the basis of protected attributes is prohibited | The do-not-use list, and the tests that show it holds |
The human-in-the-loop requirement also makes the deploying organisation responsible for the accuracy and appropriateness of the advice or decision. And in its discrimination research, Anthropic states that it does not endorse or permit using language models to make automated decisions in the high-risk use cases it studied. For an architect, that settles the shape of the system: Claude recommends, a qualified person decides.
Here is Pernille's notice for applicants. Look at how its four sentences cover disclosure, scope, documentation and appeal.
We use an AI system, Claude, to compare your CV and answers with the published requirements for this role.
It does not see your name, age, address or photo, and it does not make decisions: a Kestrelmoor recruiter reviews every application and decides who is shortlisted.
We test the system regularly to check that it treats applicants alike, and we keep a written description of how it works and what it cannot do.
If you think something in your application was missed or misread, use the "Ask for a review" link on your application page within 30 days, and a recruiter who did not handle it will look again.
Documentation is the part candidates never read and everyone else relies on. Anthropic's Transparency Hub publishes a model report for each recent Claude model, covering its acceptable uses, training data and testing results, and some of its safety summaries report bias measures such as political even-handedness. That describes the model, not your criteria or your applicants. So Kestrelmoor writes a system card of its own. It records the purpose, what the model sees and what is masked, the criteria per role and who approved them, the latest test results, known limits, the owner and the next review date. Track appeals too, because the rate at which appeals overturn outcomes, per group, is itself a fairness signal.
5.5.6 The exam traps
Every trap below makes a system look fair without measuring whether it is.
- ✗ Removing names, ages and photos, then declaring the system fair. ✓ Mask them in code, then test. Proxies such as a career gap, a school or a postcode survive masking, and only counterfactual and by-group tests show whether they still steer outcomes.
- ✗ Accepting a strong aggregate score as evidence of fairness. ✓ Break results down by group. A system at 91% overall can fail one group badly, and the aggregate will never show it.
- ✗ Treating "do not discriminate" in the prompt as the control. ✓ Keep the instruction, then prove its effect with a counterfactual test on your own cases and watch outcomes in production.
- ✗ Building criteria, examples or test labels from past decisions. ✓ Derive criteria from the job. Past decisions repeat past preferences, and a test set labelled with them scores the old bias as accuracy.
- ✗ Asking Claude whether it was biased, or showing its account of its own reasoning as the explanation. ✓ Explain each outcome with the criteria and the quoted evidence, and measure bias with tests. A self-assessment is not a measurement.
- ✗ Dropping the AI notice because a person approves each decision, or wording it vaguely. ✓ Say plainly that AI assists, what it does and does not decide, and how to ask for a human review.
Four fixes that feel like fairness, one that is
5.5.7 Put it together: test a screening prompt for bias
You now have the whole design: the six routes bias takes, the three tests that find it, the layered mitigations, and what people must be told. To make it stick, run a counterfactual test yourself, plant a biased criterion, and watch the test catch it.
Fairness work does not end at launch, and the next domain carries it on. Documentation (6.4) is where the system card lives once the engagement ends, and lifecycle support (6.5) turns the monthly by-group report and the appeal rate into standing reviews with an owner. Two neighbours own what this lesson only named: human-in-the-loop validation (5.3) designs the recruiter's review step, and regulatory compliance (5.4) settles which laws govern the group data, the notice and the records.
Key takeaways
- ✓ Fairness belongs to the whole system: bias enters through the model, the prompt and criteria, few-shot examples, retrieved text, historical data and the evaluation set.
- ✓ A counterfactual test changes one attribute or proxy and nothing else, with a no-change control to separate bias from run-to-run noise.
- ✓ Break live outcomes down by group, using group data the model never sees, because an aggregate figure can hide a group the system fails.
- ✓ Review the criteria themselves: a consistently applied criterion that is not about the job is still unfair.
- ✓ Layer the mitigations: job-related criteria, masking in code, an evidence-based rubric, a person on every consequential decision in both directions, and monitoring by group; a prompt instruction against bias is an extra layer, not the control.
- ✓ Tell people AI is involved and what it decides, document the system and its limits, and offer an appeal; for high-risk uses such as resume screening, Anthropic's Usage Policy requires qualified human review and disclosure.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
27 CCAR-P questions on Domain 5, free
Every question in the bank is tagged to a domain, so you can drill 27 questions on Governance, Safety & Risk Management alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 5 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.