Home › Study guides › CCAR-P › Domain 7 › Lesson 7.2
CCAR-P · Domain 7 · 7% of the exam · Lesson 7.2 · 22 min read
Improving developer workflows with Claude Code, and measuring the effect
Where AI coding help pays off, the Claude Code workflows that keep quality high, and how to measure the effect by cycle time and failure rate, not lines.
Written against objective 7.2 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
7.2.1 Why handing out seats rarely makes a team faster
Orrisfield is a digital-health company whose patient app lets about 400,000 patients of partner clinics book appointments, read laboratory results, message their care team and get medication reminders. Wanjiru runs the app's engineering group: 26 engineers in three squads, one each for iOS, Android and the backend services behind both apps. Two problems top every retrospective. Test coverage on the results and booking services has slipped release after release. And a pull request now waits about two and a half days for its first review, because only a few senior engineers are trusted with code that shows clinical data.
The company has just bought Claude Code for every engineer. Claude Code is Anthropic's agentic coding tool: it reads a repository, edits files and runs commands on a developer's behalf. The plan on the table is the usual one: hand out seats, give a launch talk, watch the usage dashboard. Aurelio, the architect asked to make it pay off, has seen how that ends. Developers write code faster, so more and larger pull requests join the same review queue. The dashboard fills with accepted lines, and nobody can say whether patients get fixes sooner or whether more bugs reach them.
The tool makes writing code faster, but a team's delivery is usually limited by understanding, verification and review. So the thing to design is not the seat but the AI-assisted workflow: a defined way of using Claude at the stage that holds work up. Each one has a gate that keeps people in charge of what merges, and a measure that shows whether it helped.
Seats alone, or designed workflows
Seats alone
Designed workflows
7.2.2 Where AI help pays off in the lifecycle
Here is the belief that trips teams up: a coding agent is a faster way to write features, so point it at the backlog. Yet writing code is often the shortest stage of a change. Before it comes understanding the code you are about to touch; after it come tests, review and release. Claude helps at each of those stages, but the payoff differs, and so does the check a person must keep.
| Stage | Where Claude helps | What a person still owns |
|---|---|---|
| Understanding unfamiliar code | Answers the questions you would ask a senior engineer; traces a flow end to end, such as a blood-test result from the laboratory feed to the screen | Checking the answer against the code before acting on it |
| Writing tests | Finds functions no test covers, follows the project's existing test patterns, proposes edge cases such as boundaries and error paths | Deciding what each test must assert, and reviewing that it does |
| Refactoring and migrations | Repetitive, pattern-based change across many files, in small testable steps | Slicing the work; the existing tests pass before and after each slice |
| Code review | A first pass in a fresh context for logic errors, missed edge cases and regressions | Approval: the reviewer and the author own what merges |
| Documentation | Drafts comments, READMEs and API references from the code | Accuracy, and the team's standards |
Read the right-hand column as part of the design: every row ends with a person. Understanding code carries the least risk, because nothing changes, and Anthropic's best-practice guide calls it an effective onboarding workflow that also eases the load on other engineers. Tests and migrations pay off because the work is repetitive and checkable; review and documentation pay off as first drafts that a person finishes.
The requirement that decides where to start is the team's constraint. Picture a road with one narrow bridge: widening the road before the bridge only lengthens the queue at the bridge. At Orrisfield the bridge is review, with the test gap close behind it. So Aurelio starts with two workflows: closing the test gap on the results and booking services, and preparing pull requests so they are quicker to review. Faster feature writing waits until the queue shrinks.
7.2.3 A session that ends in evidence: explore, plan, test, commit
The first failure happens inside a single session. A developer types "fix the reminder bug", Claude edits three files and reports that it is done, and the developer pushes. Anthropic's guide names the weakness: Claude stops when the work looks done, and without a check it can run, "looks done" is the only signal it has. The developer becomes the only test.
The guide's recommended session has four phases: explore, plan, implement, commit. Plan mode makes the first two safe: Claude reads files and answers questions but edits nothing until you approve its plan. You cycle to it with Shift+Tab or start a session with claude --permission-mode plan. Exploring and planning first stops Claude from solving the wrong problem well. The guide is just as clear about when to skip it: if you could describe the diff in one sentence, such as renaming a variable, ask for the change directly.
Test-first puts the gate in place before the code exists. You or Claude write the tests that define "fixed", you confirm they fail and review what they assert, and only then does Claude implement until they pass. The docs offer two versions: write a failing test that reproduces a bug and then fix it, or have one Claude session write the tests and another write the code that passes them. Either way, the check cannot quietly bend to fit whatever the code does.
One change, from question to commit
Here is the workflow on one of Orrisfield's bugs: on the night the clocks went back, Android patients with a 1:30 a.m. dose got the reminder twice. The developer sends the three steps below, approving the plan and checking each failing test against the clinical rule before she commits it. Look at step 3: the tests are committed before the fix, so any edit to them shows in the diff, and Claude is told to stop rather than bend them.
Step 1, in plan mode: trace how a reminder is scheduled, from the medication plan to the Android alarm, and list every place local time is converted. Propose a plan.
Step 2, after I approve the plan: write the tests first. Cover a reminder time that occurs twice when the clocks go back, one that does not exist when they go forward, and a patient who changes time zone between two doses. Each dose fires exactly once, at the planned local time. Run them and show me the failures.
Step 3, after I have reviewed and committed the tests: fix the scheduler until the whole reminder suite passes. Do not edit the reviewed tests; if one looks wrong, stop and tell me why.
7.2.4 Scaling out: parallel sessions, headless runs and pull-request automation
Once one session works, Claude Code offers three ways to scale beyond one developer at one prompt, and each changes who is watching the run.
Parallel sessions in git worktrees. A worktree is a separate checkout of the same repository on its own branch. Running claude --worktree booking-tests in one terminal and a second named session in another means their edits never collide. It suits independent tasks, such as one session writing tests for the iOS booking screens while another fixes a crash. The cost is human, not technical: every parallel session produces another change to review. At Orrisfield, parallel sessions come after the review workflow, not before.
Headless runs. claude -p "<prompt>" runs Claude Code with no interactive session, for scripts and continuous integration (CI). It returns a non-zero exit code when the run fails, can print JSON with --output-format json, and takes a list of pre-approved tools with --allowedTools. In an unattended run, anything else that would need a person's approval is denied. Aurelio uses one for a nightly report of public functions in the results service that no test covers.
Pull-request automation. The Claude Code GitHub Action, anthropics/claude-code-action@v1, runs Claude Code inside a repository's workflows. With no prompt input it answers when someone mentions @claude in an issue or pull request; with one, it runs on any event, such as a pull request opening or a nightly schedule. On issue and pull-request events, only users with write access can trigger it. Organisations on a Team or Enterprise plan can instead switch on Code Review, a managed service in research preview. Several agents review each pull request and post findings as inline comments tagged by severity.
| Option | When it wins | What it costs |
|---|---|---|
| Interactive session in plan mode | Ambiguous or unfamiliar work that a developer steers | The developer's attention for the whole task |
| Parallel sessions in worktrees | Independent tasks in one repository | Every output still needs a reviewer; each checkout needs its own setup |
Headless claude -p in a script or CI |
A repeatable, well-specified job, such as a nightly test-gap report | An exact prompt and tool list, since nobody can answer a question mid-run |
| GitHub Action on repository events | Team-wide automation whose workflow file is reviewed like code | GitHub Actions minutes plus API tokens, and a workflow to maintain |
| Managed Code Review | A consistent first-pass review on every pull request, with no workflow file | A Team or Enterprise plan without Zero Data Retention, research-preview status, cost that grows with pull-request size |
Orrisfield picks the Action because it keeps the prompt, model and triggers under the team's control, in a file reviewed like any other code. Look at the trigger, which fires on every new or updated pull request, and at --comment, which posts findings on the pull request instead of only in the run log. The claude_args line must name the inline-comment tool, because the action starts that tool only when this line names it.
on:
pull_request:
types: [opened, synchronize, ready_for_review, reopened] # every new or updated PR
jobs:
review:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
issues: read
id-token: write
steps:
- uses: actions/checkout@v6
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }} # a secret, never in the file
plugin_marketplaces: "https://github.com/anthropics/claude-code.git"
plugins: "code-review@claude-code-plugins"
prompt: "/code-review:code-review --comment ${{ github.repository }}/pull/${{ github.event.pull_request.number }}" # post on the PR
claude_args: '--allowedTools "mcp__github_inline_comment__create_inline_comment"'
7.2.5 Keeping quality: tests gate, people approve, changes stay small
A month in, a tempting sentence starts to circulate: Claude wrote it, Claude reviewed it, the tests are green, so it must be fine. Each part can be true while the change is wrong. Tests written in the same session can assert whatever the code happens to do. A reviewer model can find no bug in a change that misses the requirement. And a person facing an 1,800-line pull request skims it, trusting both.
Three controls keep quality, and none of them is new. The tests are the gate. CI runs the suite on every pull request, and branch protection blocks the merge while it fails. Tests that Claude wrote are reviewed for what they assert, and any change that weakens an existing test gets at least as much scrutiny as a change to production code. The guide's rule is short: if you can't verify it, don't ship it.
Changes stay small. A reviewer can judge one concern; nobody can judge sixty files at once. The session's plan slices the work, migrations move one module at a time (the docs advise refactoring in small, testable increments), and each slice becomes its own pull request.
People approve. The developer who opens a pull request owns it, whoever typed the code, and the docs say to review Claude's changes before merging. AI review is a first pass that finds and ranks problems; approval stays with a person. Code Review is built that way: its findings never approve or block a pull request. At Orrisfield, changes to how results or reminders are shown also need one of the named clinical-path reviewers, and test data uses synthetic patients only.
The path every AI-written change takes to main
7.2.6 Measuring the effect: outcomes, not output
Three months in, the board will ask whether Claude Code was worth it, and the easiest number to show is the wrong one. The analytics dashboard for Team and Enterprise plans reports lines of code accepted, suggestion accept rate and daily active users. Its GitHub integration adds the pull requests and lines written with Claude Code's help. The docs suggest reading these beside DORA metrics (from the DevOps Research and Assessment programme), such as lead time for changes and change failure rate. Lines describe activity, not value: the count ignores lines deleted later, a verbose change scores higher than a careful one, and every extra line is one more to review.
So Aurelio measures what the workflows were meant to change, plus a guardrail for what they might break.
| Metric | What it tells you | How it misleads on its own |
|---|---|---|
| Lines accepted, accept rate | Whether people use the tool | Rises with verbose code; says nothing about delivery or defects |
| Cycle time (pull request opened to merged) | How fast a change gets through | Falls when review is skipped, so read it beside failure rate |
| Review time (opened to first review, and to approval) | How long work waits for people | Falls with rubber-stamping, so the same pairing applies |
| Change failure rate | The share of releases that need a hotfix or rollback | Lags, and needs enough releases to mean anything |
| Coverage of the target modules | Whether the test gap is closing | Counts lines that ran, not whether any test checks the rule |
Method matters as much as the metrics. Take a baseline from the weeks before rollout, using exactly the definitions you will use afterwards; the pull-request history usually holds it. Roll out squad by squad, so a squad not yet onboarded acts as a comparison group for seasonal swings, such as the autumn surge in vaccination bookings. A short developer survey adds what the numbers miss, such as friction and abandoned tasks.
Adoption belongs in the report as a leading indicator: it shows the workflows are in use, not that they help. Adoption is designed, not mandated: a daily-use quota inflates the dashboard and changes nothing else. Orrisfield starts with the backend squad as a pilot, then names champions, engineers who share the prompts that worked and help colleagues get started; the docs suggest using the dashboard's leaderboard to find them. The workflows ship as shared commands committed to each repository (project skills such as /add-tests for the test-first routine), so everyone runs the same version.
MEASUREMENT PLAN, AI-assisted workflows, patient app. Owner: Wanjiru. Reviewed at the end of each quarter.
Baseline: the 8 weeks before rollout, all three squads, the definitions below.
Rollout: backend squad first; iOS and Android follow four weeks later and serve as the comparison group until then.
Primary: median wait for first review (today 2.5 days) and median cycle time; coverage of the results and booking services.
Guardrail: change failure rate must not rise. A rise pauses the rollout until the cause is found.
Context only, never reported as success: lines accepted, accept rate, active users.
7.2.7 The exam traps
Every trap here mistakes activity for improvement or hands the gate to the tool. The fix is usually to change the workflow or the measure, not the model.
- ✗ Measuring success by lines accepted or suggestion accept rate. ✓ Measure cycle time, review time and change failure rate against a baseline. Usage counts show adoption, and they rise with verbose code.
- ✗ Letting the AI review approve or merge pull requests to clear the queue. ✓ Keep AI review as a first pass. A person who owns the change approves it, and the suite blocks the merge.
- ✗ Having Claude write the code and its tests together, then trusting green. ✓ Write and review the tests first, or a failing test that reproduces the bug. Tests written from the code certify whatever the code does.
- ✗ Accepting large multi-module pull requests because Claude produced them quickly. ✓ Slice the work into small, single-purpose pull requests a reviewer can actually judge.
- ✗ Pointing Claude at feature output while review is the bottleneck. ✓ Aim it at the constraint first, such as tests and review preparation. More code into a blocked queue lengthens cycle time.
- ✗ Driving adoption with a usage mandate and personal copies of prompts. ✓ Start with a pilot, name champions, and ship the workflows as shared commands in the repository.
Four numbers that prove nothing, one that does
7.2.8 Put it together: design a gated workflow and measure it
You now have the whole design: aim at the constraint, end every session in evidence, scale with care, keep the gates human and measure outcomes. The quickest way to make it stick is to watch a test written from the code certify a bug.
Debugging and operational support (7.3) applies the same discipline to incidents: give Claude the evidence, let it propose, and confirm every hypothesis before anyone acts on it.
Key takeaways
- ✓ AI help speeds a team up only where it is slow, so aim it at the bottleneck, often tests and review, before feature output.
- ✓ Claude helps across the lifecycle (understanding code, tests, refactoring and migrations, review, documentation), and a person owns the final check at every stage.
- ✓ A good session explores and plans in plan mode, writes reviewed tests before the code, and ends with evidence that the tests ran and passed.
- ✓ Worktrees, headless
claude -p, the GitHub Action and managed Code Review scale the workflow; choose by who watches the run, and plan the review capacity it needs. - ✓ Quality gates stay human: a suite that blocks the merge, small single-purpose pull requests, and a person who approves; AI review is a first pass.
- ✓ Measure cycle time, review time and change failure rate against a baseline and a comparison group; lines accepted show adoption, which grows through pilots, champions and shared commands.
Check your understanding
4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
12 CCAR-P questions on Domain 7, free
Every question in the bank is tagged to a domain, so you can drill 12 questions on Developer Productivity & Operational Enablement alone, or sit the full 63-question timed simulator.
Open the CCAR-P question bank → Back to Domain 7 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.