Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-F › Domain 3 › Lesson 3.6

CCAR-F · Domain 3 · 20% of the exam · Lesson 3.6 · 23 min read

Claude Code in CI/CD pipelines

Run Claude Code in a CI pipeline: -p so jobs never hang, --json-schema for PR comments, prior findings and CLAUDE.md for context, a separate reviewer.

Written against task statement 3.6 of the official CCAR-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

3.6.1 Why a pipeline job is not a terminal

Think about the difference between asking a colleague for help across a desk and leaving a note for the night shift. Across the desk you can be vague, because they can ask what you meant. The note has to be complete: what to do, where things are, what form the answer should take and where to leave it. Nobody will be awake to answer a question at three in the morning.

Software teams have their own night shift. A pull request (PR) is a proposed code change waiting to join the main code. When a developer opens one, a continuous integration (CI) pipeline starts: a borrowed machine called a runner fetches the code and works through a list of automated jobs, such as running the tests. (The CD in CI/CD is continuous deployment: the jobs that ship the code once it passes.) No human is attached to any job: no keyboard, no screen, no one to click "allow". Each job must start, do its one thing, leave its result where a script can read it, and exit.

Claude Code, Anthropic's command-line coding agent, was built first for conversation. Type claude and it opens an interactive session: it answers, asks before running some commands, and waits for your next message until you leave. That is right at a desk and wrong on a runner. So Claude Code has a second way to run, non-interactive mode (also called print mode or headless mode). It takes one prompt, does the work, prints the result and exits.

We will follow one pipeline the whole way, belonging to an online shop's development team. Its main job is an automated security review. On every pull request it looks for vulnerabilities in the changed code and posts each finding as an inline comment, pinned to the exact line. A second job writes tests for the changed code. Along the way the pipeline will hang, produce prose no script can read, repeat itself, suggest tests that already exist and grade its own homework. Each failure has one fix.

3.6.2 The job that never finished: interactive versus headless

Here is the first version of the review job, and the first thing that goes wrong. The job runs claude "Review this pull request for security vulnerabilities". It starts, and then nothing happens. Ten minutes later the runner kills it for exceeding its time limit. No review, no comments, one red cross on the pull request, and a developer with no idea why.

claude with a prompt starts an interactive session with that prompt as its first message. It answers, then waits for your next message, and it may stop earlier to ask permission before running a command. On the runner nobody will ever type that message, so it waits forever. From the outside that looks like a hang. It is not a crash; it is Claude Code behaving exactly as it does at your desk, in an empty room.

The fix is one flag. -p, long form --print, runs Claude Code non-interactively: read the prompt, do the work, print the result to standard output (the text stream a script can capture), and exit. The command becomes claude -p "Review this pull request for security vulnerabilities", and the job now ends when the review ends. Like any command-line tool, it also returns an exit code, 0 on success and non-zero on failure, so the pipeline can tell a finished review from a broken one.

Interactive versus headless

Interactive claude

Opens a sessionexpects a keyboard
Waits for the next message
Job never endsthe runner times out

Headless claude -p

Reads the promptargument or piped input
Does the work, prints the result
Exits with a status code
An interactive session waits for a person; a headless run takes its prompt from the command line or from piped input, prints a result and exits.

Two habits go with -p. First, print mode reads standard input, so the job can pipe the diff (the lines the pull request adds and removes) straight in: gh pr diff 123 | claude -p "...", where gh is GitHub's command-line tool. Claude then has the changed code in front of it without needing permission to run a git command. Second, nobody is there to approve a tool at run time, so in print mode a tool call that would need approval is refused. The job pre-approves the tools it needs with --allowedTools. Neither habit replaces -p: an interactive session with every tool pre-approved still waits for your next message.

3.6.3 Findings a script can post: --output-format json and --json-schema

With -p in place, the job finishes and prints a good review: three clear paragraphs. One says the new query in api/orders.py pastes the customer's input straight into the database query, so an attacker can rewrite it (an SQL injection) on line 42. The next step is meant to post each finding as an inline comment. For that it needs, for every finding, a file path, a line number, a severity and a message. What it has is prose. Someone would have to read it and copy the numbers out, and that someone is the human the pipeline was built to remove.

Prose is fine for a person and useless for a program. A script needs a machine-parseable shape: the same fields, in the same place, every time. Claude Code provides this in two layers, both for print mode only. The first is --output-format json. Instead of plain text, the run prints one JSON object (JavaScript Object Notation, the standard text format for structured data). The answer sits in a field called result, next to metadata such as the session ID, the cost and usage figures. The script now has a reliable envelope, but the review inside it is still prose.

The second layer is --json-schema. A JSON Schema describes the shape you want: which fields, of which type, and which ones are required. You pass one with this flag, together with --output-format json. Claude Code checks the final answer against it, asks the model to try again if it does not match, and puts the validated object in a field called structured_output. The model can still read files and run tools during the review; the schema applies only to the answer it gives at the end. Think of a form with numbered boxes handed to an expert witness. They may investigate however they like, but the report comes back with every box filled in and nothing scribbled in the margins.

What you pass What the script receives Where the review sits
-p alone (--output-format text, the default) Plain text, as a person would read it Somewhere in the prose
--output-format json One JSON object: the answer plus session ID, cost and usage In result, still as prose
--output-format json plus --json-schema The same object, plus the answer checked against your schema In structured_output, as fixed fields
--output-format stream-json One JSON object per line, as events happen In the final result line; built for live progress, not for this job

Memorise the middle two rows and which field holds what: result for the text, structured_output for the schema-checked object. Recognise text and stream-json as the other two values of --output-format.

Here is the job's review step. --append-system-prompt adds a reviewer role to Claude Code's built-in instructions for this run. The lines that matter are the two output flags and the jq call (jq is a small tool for pulling fields out of JSON) that extracts the findings from structured_output.

gh pr diff "$PR_NUMBER" | claude -p \
  --append-system-prompt "You are a security reviewer. Report only real vulnerabilities." \
  --output-format json \
  --json-schema "$(cat review-schema.json)" \
  | jq '.structured_output.findings' > findings.json   # one object per finding

And the schema in review-schema.json. It says a review is a list of findings, and every finding must carry the four fields the comment step needs:

{"type": "object",
 "properties": {"findings": {"type": "array", "items": {
   "type": "object",
   "properties": {"file": {"type": "string"}, "line": {"type": "integer"},
                  "severity": {"type": "string", "enum": ["high", "medium", "low"]},
                  "message": {"type": "string"}},
   "required": ["file", "line", "severity", "message"]}}},
 "required": ["findings"]}

From here on, no model is involved. An ordinary script reads findings.json and, for each entry, calls the code host's API (application programming interface, the way one program asks another to do something) to post a comment at file and line. Claude produces the findings; your pipeline posts them.

One run of the review job

CHECKOUTthe PR branch and its diff
REVIEWclaude -p with --json-schema
PARSEread structured_output
POSTone inline comment per finding
Every step after the model call is plain scripting, which is only possible because the findings come back as validated JSON.

3.6.4 The context a fresh run does not have

The job works. Then a developer pushes two more commits (saved sets of changes) to fix the SQL injection, the pipeline runs again, and the pull request fills up with the same comments a second time. Some point at the line the developer just fixed. Now the developer has to read every comment to work out which are new, which is worse than having no bot at all. The team's first instinct is to add "avoid duplicate comments" to the prompt. It changes nothing, and it could not.

Each claude -p run starts a new session with no memory of the previous one. The second run never saw the first run's findings, so every issue it finds looks new to it. Asking it not to repeat itself is like asking the night shift not to redo yesterday's work without saying what yesterday's work was. The information has to be in the note. So the job keeps each run's findings.json (or reads back the comments it posted last time). On the next run it puts those prior findings in the prompt with an instruction: report only issues that are new or still unaddressed in this diff.

The second run, with and without memory

Run 2, fresh context

Sees the new diff only
Re-finds the same issues
Posts duplicate comments

Run 2, prior findings in context

Sees the new diff
Sees last run's findingssupplied by the job
Reports only new or still-open issues
A new run knows nothing about the last one; the pipeline has to hand it the earlier findings and say what to do with them.

The test-generation job has the same disease in a different coat. It keeps suggesting a test for "order with an empty basket", a test that has existed in tests/test_orders.py for two years. The model is not careless. It was given the diff and asked for tests, and nothing told it which tests already exist. The fix has the same shape: the job puts the existing test files for the changed code in context. Now the model can see which scenarios are covered and suggest only the missing ones.

That is the pattern: a CI run cannot know what it was not given. Duplicate comments and duplicate tests are both cured by supplying what the run would otherwise have to remember. A stronger prompt without that context is a request the model has no way to honour.

3.6.5 CLAUDE.md: the standing instructions for every run

Prior findings and existing tests change with every pull request, so the job supplies them per run. Other things are true for every run: how this team writes tests, what makes a test worth having, which ready-made test helpers exist, what the reviewer should and should not flag. Pasting those into every job's prompt means a copy in every pipeline file, and the copies drift apart. The team already has a better place for them.

CLAUDE.md is a plain-text file of standing instructions that Claude Code reads at the start of every session. The project's copy lives at the root of the repository (the project's shared code store), or at .claude/CLAUDE.md. It is committed with the code, so every developer and every job gets the same version. Unless told otherwise, a claude -p run loads the same context an interactive session would: the runner fetches the code, and CLAUDE.md comes with it. Think of the laminated card pinned above the night-shift desk. It never changes per task, and it is read before every task.

For the test-generation job, three things belong there. The testing standards: framework, naming, where tests live. The criteria for a valuable test: what a test must check to earn its place. And the available fixtures: the ready-made test data and setup the test suite already provides. Before the team wrote these down, the generated tests were plentiful and cheap. There was a test for every trivial one-line function, home-made fake objects for things the fixtures already supplied, and three tests asserting the same thing. Afterwards the output shrank and improved, because the model knew what "good" meant here and what already existed.

The review job gains the same way from a short review-criteria section: what counts as a finding and what to ignore.

# Testing standards
- pytest; tests live in tests/, one file per module, named test_<module>.py
- Use the fixtures in tests/conftest.py: `db_session`, `sample_order`, `auth_client`
- A valuable test checks a behaviour or an edge case, not a one-line getter
- Do not build fake objects (mocks) for anything a fixture already provides

# Review criteria (security job)
- Report injection, broken auth, secrets in code, unsafe deserialisation
- Do not report style, naming or missing docstrings

Keep it short and concrete. The docs suggest under 200 lines, and instructions you can check ("use the db_session fixture") beat ones you cannot ("write good tests"). The docs for running Claude Code in GitHub Actions, GitHub's pipeline service, say the same: review criteria go in the root CLAUDE.md, kept concise because Claude reads it on every run. Remember what it is: context the model reads and tries to follow, not an enforced rule. Anything that must happen every time belongs in a hook (a script Claude Code runs at a fixed moment) or a permission rule.

3.6.6 Why the author should not be the reviewer

The last design decision is the subtle one. The team extends the pipeline: when a ticket gets a certain label, a Claude Code job implements the change and opens the pull request. It seems efficient to let that same run review its work before pushing, since it already understands the change. So the job's prompt ends with "now review your changes for security issues". The reviews always come back clean. Weeks later a human finds an injection bug that went straight through.

The model is not bad at review; the trouble is that this reviewer knows why every line is the way it is. The session that wrote the code holds the plan, the assumptions and the reason its shortcut in a database query felt safe. Asked to review, it reads the code through that reasoning and confirms what it already believes, like proofreading your own essay: you read what you meant, not what you wrote. The fix is session context isolation. Claude Code's docs put it plainly: a fresh context improves code review because Claude is not biased towards code it just wrote.

In a pipeline, isolation is cheap. The review becomes its own claude -p invocation, an independent instance that receives the diff and the review criteria and nothing else. Two shortcuts quietly undo it. Running the review with --continue, which reopens the most recent conversation, hands the reviewer the author's whole session back. Pasting the implementing run's explanation into the review prompt does the same in a smaller dose. And telling the same session to "review critically" does not help: the instruction changes the question, not what the session is holding.

Who reviews the change

Same session

Holds the plan it followed
Shares its own assumptions
Confirms what it already believes

Independent instance

Sees the diff and the criteria
No memory of why the code is that way
Questions every line on its merits
The session that wrote the code carries the reasoning that produced it; an independent instance sees only the diff and the criteria and judges the result on its own terms.

This does not contradict the advice to give a run more context. Prior findings give the reviewer facts about the pull request; the author's session gives it the author's beliefs. The first makes the review sharper; the second makes it agree with itself.

3.6.7 The exam traps

Every mistake here treats a CI run like a person at a desk. It assumes the run will wait, that its prose will be read, that it remembers, that it knows the house rules, or that it can judge its own work.

  • ✗ Fixing a hanging job with a longer timeout or a made-up switch such as --headless or --non-interactive. ✓ Add -p (--print), the documented flag. The job is not slow; it is waiting for a keyboard, and a longer timeout only decides when it fails.
  • ✗ Asking for "file:line - message" in the prompt and parsing the prose with patterns. ✓ Use --output-format json with --json-schema and read structured_output. Prose drifts from run to run; a schema does not.
  • ✗ Asking Claude to "avoid duplicate comments" on a re-run. ✓ Include the prior findings in context and ask for only new or still-unaddressed issues. A fresh run cannot avoid repeating what it never saw.
  • ✗ Telling the test generator to "avoid tests that already exist". ✓ Provide the existing test files in context. Same disease, same cure.
  • ✗ Copying testing standards into every job's prompt, or leaving them out. ✓ Document testing standards, valuable-test criteria and fixtures in CLAUDE.md, which every run loads.
  • ✗ Letting the session that wrote the code review it, in the same run or through --continue. ✓ Review in a separate claude -p invocation that sees only the diff and the criteria.

3.6.8 Put it together: wire the review job and break it

You now have the whole job: a headless run that cannot hang, and output a script can post. Per-run context covers what changes, CLAUDE.md covers what does not, and the reviewer did not write the code. The fastest way to make the flags stick is to run them, then remove them one at a time and watch each failure.

The structured-output half of this lesson continues in Domain 4: schemas that enforce JSON output (4.3) and review criteria explicit enough to keep false positives rare (4.1). Multi-instance and multi-pass review (4.6) builds on the independent reviewer you met here. The context half continues in Domain 5, where what a run is given, and what it must not be given, becomes the whole subject.

Key takeaways

  • ✓ A CI runner has no keyboard and no person, so a job must start, do one thing, leave a result a script can read, and exit.
  • ✓ -p (--print) runs Claude Code non-interactively; without it the job opens an interactive session and hangs until the runner's time limit kills it.
  • ✓ --output-format takes text, json or stream-json; with json, --json-schema checks the final answer against your schema and returns it in structured_output, ready to post as inline comments.
  • ✓ A fresh run remembers nothing, so a re-run needs the prior findings and an instruction to report only new or still-unaddressed issues, and test generation needs the existing test files.
  • ✓ CLAUDE.md loads at the start of every run, so testing standards, valuable-test criteria, fixtures and review criteria belong there.
  • ✓ The session that wrote a change is a weak reviewer of it, so the review runs as a separate claude -p invocation that sees only the diff and the criteria.

Check your understanding

4 questions written for this lesson, then one from the CCAR-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

72 CCAR-F questions on Domain 3, free

Every question in the bank is tagged to a domain, so you can drill 72 questions on Claude Code Configuration & Workflows alone, or sit the full 60-question timed simulator.

Open the CCAR-F question bank → Back to Domain 3 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources