Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 5 › Lesson 5.5

CCAR-P · Domain 5 · 14% of the exam · Lesson 5.5 · 23 min read

Bias, fairness and transparency when Claude helps decide about people

Where bias enters a Claude solution, how to test for it with counterfactual pairs and group breakdowns, which mitigations hold, and what people must be told.

Written against objective 5.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.5.1 Why a consistent shortlist can still be unfair

Kestrelmoor Staffing fills warehouse and office jobs at about forty client sites: night-shift pickers, forklift drivers, receptionists and payroll clerks. Around 3,000 applications arrive each week, and eleven recruiters cannot read them all before the best candidates accept other offers. Nnamdi, the operations director, wants Claude to read each CV and the answers to the application questions, then hand recruiters a shortlist for every role. Pernille, the solution architect on the engagement, is confident that Claude can do the reading. What worries her is something else.

Nobody at Kestrelmoor means to be unfair. The risk is that the system will be unfair anyway, quietly and at scale. A recruiter who marks down a CV for a two-year gap does it to a handful of people a week. A prompt that says "prefer continuous employment" does it to every parent, carer and person who has been ill, thousands of times, and every one of those decisions looks consistent. Consistency feels like fairness, but a system can apply the wrong rule perfectly.

Let's fix three terms first. Bias here means a systematic difference in how a system treats people that comes from something other than what the decision is about. Fairness is the property you design and test for: comparable people get comparable outcomes, and the criteria themselves are about the job. Transparency is what people are told and what is written down: that AI is involved, what it decides, where its limits are and how to challenge a result. None of the three is a setting on the model, so none of them can be handed to it.

5.5.2 Where bias gets in

Here is the belief that trips people up: bias is something the model has, so the fix is a better model. Models can carry bias. Claude learned language from human writing, which is full of associations about which jobs suit which people and what a "professional" CV sounds like. But in a system like Kestrelmoor's, the model is only one of six places where bias gets in, and the architect's own pipeline supplies the other five.

Think of a cake that comes out wrong: it might be the oven, but more often it is the recipe or the ingredients. Pernille walked Nnamdi's team through the design and found an example of each route.

Where it enters How it gets in What it looked like at Kestrelmoor
The model Tendencies learned from human writing, such as reading polished prose as competence Higher scores for fluent, idiomatic English on warehouse roles that need only basic reading
The prompt and criteria A criterion unrelated to the job, or a vague one the model fills with its own assumptions "Prefer continuous employment" and "a young, energetic team", copied from a client brief
Few-shot examples Examples teach the pattern of who was chosen, not only the output format Three "strong candidate" examples for forklift roles, all men in their twenties
The retrieval corpus Retrieved text brings its authors' assumptions into the context Old job adverts retrieved as role descriptions, one asking for "strong lads"
Historical data Past decisions used as labels or targets repeat past preferences A plan to "shortlist people like the placements clients kept"
The evaluation set Labels copied from past decisions score the old bias as accuracy, and missing groups hide failures Last year's recruiter picks as the answer key, with almost no applicants over 55

Memorise the six routes; in a scenario, the symptom usually points to one of them. The last row deserves extra attention, because the evaluation set decides whether you ever see the other five. If your answer key is last year's shortlists, a system that copies last year's preferences scores well, and the test reports success.

5.5.3 Testing for fairness: change one thing, then count by group

How would Pernille know whether the shortlist is fair? Asking Claude is not a test: a model's account of its own reasoning is not a measurement of its behaviour. Anthropic's evaluation guide makes the point that even hazy qualities such as ethics and safety can be turned into measurable criteria. For fairness, three tests do the measuring, and each catches something the others miss.

The first is the counterfactual test. Take a CV, make a copy that differs in exactly one attribute, run both through the same prompt, and compare the results. It works like a wine tasting with swapped labels: pour the same wine into two glasses, change only the label, and if the ratings differ, the label did it.

The attribute can be a protected one, such as a name that signals gender or ethnicity. It can also be a proxy, a detail that stands in for a protected one, such as a graduation year for age or a career break for caring. Claude's outputs vary a little from run to run, so the test needs a no-change control, the identical CV run several times. A gap counts only when it is bigger than the control's spread.

This is the method Anthropic used in its 2023 research on discrimination in model decisions. The researchers generated 70 decision scenarios, such as whether to offer a job or pay out an insurance claim, with slots for age, race and gender, and varied only those. With no mitigation, Claude 2.0 favoured some groups and disfavoured others in some settings. Plain-language instructions reduced both, and two nearly eliminated them in those scenarios: saying that discrimination is illegal, and asking the model to answer as if no demographic information had been given. Borrow the method, not the result. The study measured one model on its own scenarios; your criteria, your applicants and today's model are untested until you test them.

In Pernille's harness, MODEL is the model Kestrelmoor plans to ship and RUBRIC is the scoring rubric. The call uses structured outputs: given a Pydantic class, messages.parse returns every reply as a valid score with its evidence, so no run is lost to a parsing error. Look at the control, which changes nothing, and the variant, which changes only the name.

from pydantic import BaseModel
import anthropic

client = anthropic.Anthropic()

class Assessment(BaseModel):
    evidence: list[str]   # one quote per rubric criterion
    score: int            # 0 to 12, from the rubric

def score(cv: str) -> int:
    reply = client.messages.parse(model=MODEL, max_tokens=4000, system=RUBRIC,
        messages=[{"role": "user", "content": cv}], output_format=Assessment)
    return reply.parsed_output.score

for cv in BASE_CVS:  # fictional CVs with a [NAME] slot
    control = [score(cv.replace("[NAME]", NAME_A)) for _ in range(3)]  # nothing changes: the noise
    variant = [score(cv.replace("[NAME]", NAME_B)) for _ in range(3)]  # ONLY the name changes
    print(max(control) - min(control), sum(variant) / 3 - sum(control) / 3)  # spread, then gap

One CV, three versions

Control identical CV, three runs

Scores 9, 9, 6

spread of 3: the noise

Name swapped

Same CV, different name
Scores 9, 6, 9

gap of 0: inside the noise

Caring break added

Same CV plus "2021 to 2023: caring for a parent"
Scores 3, 6, 3

gap of 4: find the cause

The control shows how far scores move when nothing changes; only a gap larger than that spread is evidence of bias.

The second test runs on real traffic: results broken down by group. Pernille compares the shortlist rate, the average score and the rate of recruiter overrides for each group, using the answers applicants give on Kestrelmoor's voluntary equal-opportunities form. That data is held apart from the screening input: the model never sees it, and only the analysis job does. An aggregate hides groups, so 91% agreement with recruiters overall can be 95% for one group and 70% for another. Agree the gap that triggers an investigation with HR and legal before launch; how the group data may be collected and kept is a compliance question.

The third test is a review of the criteria themselves. The first two check that the system applies its criteria consistently; neither can say whether a criterion belongs there at all. "Prefer continuous employment" passes a name swap perfectly and still excludes carers for a reason unrelated to the job. For each criterion, ask which job requirement it measures and whether someone who fails it could still do the work. Pernille's review struck out the continuity rule and limited English on warehouse roles to reading pick lists and safety signs. She put each struck rule to the client as a choice: name the job requirement it measures, or drop it.

5.5.4 Mitigations, from the criteria outwards

Once a test finds a gap, the tempting fix is one line in the prompt: "Do not discriminate." Keep the line, since Anthropic's research found that instructions like it can cut measured discrimination sharply. But an instruction is guidance you then have to verify, and it cannot repair a criterion that is unfair by design. The mitigations that hold are layered in the order a CV meets them: fix what goes in, constrain how the judgement is made, and control what happens after.

The controls in the order a CV meets them

CRITERIAjob-related, agreed per role
MASKcode removes name, age, photo, address
RUBRICevery judgement tied to a quote
REVIEWa recruiter decides, both ways
MONITORoutcomes by group, monthly
Preventive controls on the input and the judgement come first; the person and the monitoring catch what gets past them.

Pernille's pipeline masks in code, before the call. Telling Claude to "ignore the name" leaves the name in the context; removing it means there is nothing to ignore. Where a client has a real need behind an attribute, replace the attribute with the fact the job needs. A client that asks for postcodes because night shifts start after the last bus gets a field computed in code, "can reach site 14 by 22:00: yes or no", instead of an address. Masking does not remove proxies, which is why the tests stay.

The rubric turns a vague judgement into a set of checks, each backed by evidence. Look at the last two lines of Pernille's rubric for one role: the do-not-use list, and the rule that Claude never outputs a rejection.

Role: night-shift warehouse operative, client site 14. Mark each criterion Met, Not met or Not stated, and quote the words in the application that support it. Each Met scores 3 points.
1. Shift pattern: available 22:00 to 06:00, four nights a week.
2. Manual handling: experience of, or willingness to, lift loads up to 20 kg.
3. Counterbalance forklift certificate: held, or booked before the start date.
4. Reading: can read pick lists and safety signs in English; do not reward style, spelling or fluency beyond this.
Do not use or infer: name, age, gender, ethnicity, nationality, religion, disability, family status, gaps in employment, address, photo, school names or hobbies. If something is not stated, record Not stated; never guess.
Output the four results with their quotes, then Shortlist or Refer to recruiter. Never reject: a recruiter decides every application.
Mitigation When it wins What it costs
Job-related criteria, agreed per role Always; every other control depends on them Client conversations, and some client requests refused
Masking in code before the model call The attribute is irrelevant and can be removed reliably Proxies survive; recruiters need a way back to the person
Evidence-based rubric Outcomes must be explained, audited and compared Design time, longer outputs, less room for judgement calls
Anti-bias instruction in the prompt A cheap extra layer on any decision prompt Proves nothing alone; test it like any other change
Human review before the decision is final Consequential decisions about people; required for high-risk uses Recruiter time; a reviewer who sees only the shortlist becomes a rubber stamp
Outcome monitoring by group Every live system that decides about people Group data held apart, a lawful basis for it, and an owner who acts

Memorise the first column; the costs are what you argue about when someone wants to drop a layer. The review row hides the subtlest trap. If recruiters see only Claude's shortlist, they can turn down a weak candidate Claude liked, but they never notice the strong candidate Claude left out, and wrongful exclusion is where unfairness usually hides. So every application reaches a recruiter with the quotes that make review quick, and each assessment is stored with the rubric version and model id so any outcome can be traced later.

5.5.5 Transparency, and what Anthropic requires for high-risk uses

A candidate Kestrelmoor turns down has three fair questions: was a machine involved, what did it decide, and how can I challenge it? Transparency is the design that answers them, and it has four parts. Disclose that AI is involved, in plain words. Scope it: say what the AI does and what it does not decide. Document the system and its limits. And provide an appeal route to a person who can change the result.

Anthropic's Usage Policy (the version in effect since 15 September 2025) turns part of this into requirements. It defines High-Risk Use Cases as consumer-facing uses in domains vital to public welfare and social equity. There are seven: legal; healthcare; insurance; finance; employment and housing; academic testing, accreditation and admissions; and media content generated and published automatically. The employment category names resume screening and hiring tools, so Pernille treats Kestrelmoor's shortlist as a high-risk use.

Requirement What the Usage Policy says today Kestrelmoor's design
Human in the loop (high-risk uses) For advice, recommendations or subjective decisions that directly affect people, a qualified professional in the field reviews the content or decision before it is disseminated or finalised A recruiter confirms every shortlist and every referral
Disclosure (high-risk uses) When outputs are presented directly to people, tell them AI helped produce the advice, decision or recommendation, at least at the start of each session The notice on the application form and in every outcome message
Chatbots (any use) A consumer-facing chatbot or external-facing agent tells users it is AI, not a human, at least at the start of each chat session The candidate help chat opens by saying it is an AI assistant
No discrimination (all users) Promoting discriminatory practices against people on the basis of protected attributes is prohibited The do-not-use list, and the tests that show it holds

The human-in-the-loop requirement also makes the deploying organisation responsible for the accuracy and appropriateness of the advice or decision. And in its discrimination research, Anthropic states that it does not endorse or permit using language models to make automated decisions in the high-risk use cases it studied. For an architect, that settles the shape of the system: Claude recommends, a qualified person decides.

Here is Pernille's notice for applicants. Look at how its four sentences cover disclosure, scope, documentation and appeal.

We use an AI system, Claude, to compare your CV and answers with the published requirements for this role.
It does not see your name, age, address or photo, and it does not make decisions: a Kestrelmoor recruiter reviews every application and decides who is shortlisted.
We test the system regularly to check that it treats applicants alike, and we keep a written description of how it works and what it cannot do.
If you think something in your application was missed or misread, use the "Ask for a review" link on your application page within 30 days, and a recruiter who did not handle it will look again.

Documentation is the part candidates never read and everyone else relies on. Anthropic's Transparency Hub publishes a model report for each recent Claude model, covering its acceptable uses, training data and testing results, and some of its safety summaries report bias measures such as political even-handedness. That describes the model, not your criteria or your applicants. So Kestrelmoor writes a system card of its own. It records the purpose, what the model sees and what is masked, the criteria per role and who approved them, the latest test results, known limits, the owner and the next review date. Track appeals too, because the rate at which appeals overturn outcomes, per group, is itself a fairness signal.

5.5.6 The exam traps

Every trap below makes a system look fair without measuring whether it is.

  • ✗ Removing names, ages and photos, then declaring the system fair. ✓ Mask them in code, then test. Proxies such as a career gap, a school or a postcode survive masking, and only counterfactual and by-group tests show whether they still steer outcomes.
  • ✗ Accepting a strong aggregate score as evidence of fairness. ✓ Break results down by group. A system at 91% overall can fail one group badly, and the aggregate will never show it.
  • ✗ Treating "do not discriminate" in the prompt as the control. ✓ Keep the instruction, then prove its effect with a counterfactual test on your own cases and watch outcomes in production.
  • ✗ Building criteria, examples or test labels from past decisions. ✓ Derive criteria from the job. Past decisions repeat past preferences, and a test set labelled with them scores the old bias as accuracy.
  • ✗ Asking Claude whether it was biased, or showing its account of its own reasoning as the explanation. ✓ Explain each outcome with the criteria and the quoted evidence, and measure bias with tests. A self-assessment is not a measurement.
  • ✗ Dropping the AI notice because a person approves each decision, or wording it vaguely. ✓ Say plainly that AI assists, what it does and does not decide, and how to ask for a human review.

Four fixes that feel like fairness, one that is

Delete the protected fieldsproxies remain
"Do not discriminate"an untested instruction
A bigger modelsame criteria, same examples
Ask Claude if it was fairnot a measurement
Test, then fix the causecounterfactual and by-group tests, job-related criteria, a person on the decision
Each tempting fix changes how the system looks without measuring how it behaves; the real fix tests first, then changes the cause.

5.5.7 Put it together: test a screening prompt for bias

You now have the whole design: the six routes bias takes, the three tests that find it, the layered mitigations, and what people must be told. To make it stick, run a counterfactual test yourself, plant a biased criterion, and watch the test catch it.

Fairness work does not end at launch, and the next domain carries it on. Documentation (6.4) is where the system card lives once the engagement ends, and lifecycle support (6.5) turns the monthly by-group report and the appeal rate into standing reviews with an owner. Two neighbours own what this lesson only named: human-in-the-loop validation (5.3) designs the recruiter's review step, and regulatory compliance (5.4) settles which laws govern the group data, the notice and the records.

Key takeaways

  • ✓ Fairness belongs to the whole system: bias enters through the model, the prompt and criteria, few-shot examples, retrieved text, historical data and the evaluation set.
  • ✓ A counterfactual test changes one attribute or proxy and nothing else, with a no-change control to separate bias from run-to-run noise.
  • ✓ Break live outcomes down by group, using group data the model never sees, because an aggregate figure can hide a group the system fails.
  • ✓ Review the criteria themselves: a consistently applied criterion that is not about the job is still unfair.
  • ✓ Layer the mitigations: job-related criteria, masking in code, an evidence-based rubric, a person on every consequential decision in both directions, and monitoring by group; a prompt instruction against bias is an extra layer, not the control.
  • ✓ Tell people AI is involved and what it decides, document the system and its limits, and offer an appeal; for high-risk uses such as resume screening, Anthropic's Usage Policy requires qualified human review and disclosure.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

27 CCAR-P questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 27 questions on Governance, Safety & Risk Management alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources