Home › Study guides › CCAO-F › Domain 7 › Lesson 7.2
CCAO-F · Domain 7 · 10% of the exam · Lesson 7.2 · 19 min read
Adjusting your approach from feedback and results
How to turn reviewer edits, complaints and results into the right change to a Claude workflow: find the pattern, agree what good means, test, close the loop.
Written against objective 7.2 of the official CCAO-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
7.2.1 Why the loudest feedback is a poor guide
It's Monday morning, eight weeks into a new routine. You're the social-media manager at the tourism board of a mid-sized coastal city. Each week you paste the events calendar into your Claude Project, and Claude drafts 20 posts: upcoming events, attraction spotlights and tips for visitors. Felix, the content editor, reviews the batch before it's scheduled.
Now the feedback has piled up. Engagement on event posts, meaning the likes, shares and clicks each one earns, has fallen week after week, while the other posts hold steady. The legal team has flagged two posts with claims nobody could verify, such as "the oldest working lighthouse in the country". And Felix spends about 40 minutes on every batch, mostly shortening posts and toning down what he calls "the brochure voice".
Everyone has a fix. Bronwen, who runs the events programme, says the published event posts leave out too much. A colleague suggests the most capable model, and you're tempted to rewrite the whole Project. Each fix answers one signal, usually the loudest, and could make another worse.
Feedback is any signal about how the output did, from a reviewer's edit to a number on a dashboard. Unlike diagnosing one bad output, adjusting from feedback means reading many signals over time and deciding what, if anything, to change.
From feedback to a tested change
7.2.2 Read the pattern, not the loudest comment
Here's the first question, and the one that trips people up: which feedback do you act on first? The newest comment, or the most senior person's, is tempting, but neither tells you how often a problem happens or what it costs. Start by putting the different kinds side by side.
| Signal | What it tells you | What it can't tell you alone |
|---|---|---|
| Reviewer edits | What people keep changing, and how | Whether a change was needed or just taste |
| Stakeholder comments | What one person wants or dislikes | How common the problem is |
| Error or rejection rates | How often output fails a check | Why it fails |
| Time spent fixing | What the gaps cost the team | Which gap costs the most |
| Outcome measures | Whether the output did its job: clicks, replies, bookings | Whether Claude caused the change |
Memorise the five kinds; the right-hand column is why no single one is enough. The tourism workflow has all five: Felix's edits, Bronwen's requests, the legal flags, the 40 minutes and the engagement figures.
Think of blood pressure. One high reading, taken after you ran for the bus, says little; a week of morning readings is what a doctor acts on. A single comment is a data point. A pattern, the same problem across many outputs, is a finding.
Claude is good at the sorting, as long as you keep the deciding. In this request, look at the second line, which asks for examples behind every count, and the last line, which keeps the decision with you.
The attached file holds 160 posts from the last eight weeks: Claude's draft of each one and the version Felix published.
Sort every edit Felix made into categories, such as length, tone, facts, links or hashtags. For each category, give the number of posts affected and quote two typical before-and-after pairs.
List separately every edit that changed a fact, with the post and the week.
Compare event posts with the other posts, week by week: average draft length, average published length and the most common edit.
Do not recommend changes yet. I will decide what to change after checking your counts.
Before you trust the counts, check a dozen posts by hand. Here they hold up: about 70% of Felix's edits are length and tone on event posts, which he cuts from about 85 words to 40. The drafts grew in week three, when you added Bronwen's request for full details to the instructions.
One caution: patterns are the rule for quality and taste, not for serious errors. Two unverified claims in 160 posts is a low rate, but a false claim published under the tourism board's name is a legal problem, so one is enough to act on.
7.2.3 Agree what good means before you change anything
Put the comments side by side and they disagree. Felix cuts event posts to about 40 words because, he says, short posts do better in busy feeds. Bronwen wants times, prices, age limits and the booking link in every post. Legal wants no claim that can't be checked. Every round of prompt changes so far has pleased one of them and annoyed another.
It's tempting to treat this as a wording puzzle. Resist it. Conflicting feedback usually means nobody ever agreed what a good post is, and no wording can hit two targets at once. Picture a taxi with three passengers calling out three addresses: a better driver doesn't help until the passengers agree where they're going.
So the conflict goes to a person, not to the prompt. The decision belongs to whoever is accountable for the output, here Imogen, the head of marketing. Bring her the pattern, a few examples and the three positions, stated fairly. Claude can help you lay out the conflict neutrally, but it shouldn't pick the winner, because a compromise it invents is only a guess about your organisation's priorities.
Conflicting feedback, before and after agreed criteria
No agreed criteria
Criteria set by the owner
Imogen, Felix and Bronwen agree shared acceptance criteria, the tests a post must pass before it goes out. Look at the second one, which settles the dispute by moving Bronwen's details to the events page.
Event posts: acceptance criteria, agreed by Imogen (head of marketing) with Felix and Bronwen
1. Under 45 words, with the hook in the first line.
2. Says what, where and when. Prices, age limits and access details live on the events page, and the post links to it.
3. Friendly and local, like a resident recommending something. No "unmissable", "world-class" or other brochure words.
4. Every claim about an attraction appears in the approved fact sheet; anything else is left out.
5. One call to action and one link.
7.2.4 Choose the level of change
With the criteria agreed, where do you make the change? This is where adjustments often go wrong: people change whatever is easiest to reach, usually the next prompt. A problem in one batch needs a one-off fix; a problem in every batch needs a fix that lasts.
| Level | What changes | When the feedback points here |
|---|---|---|
| The prompt | This week's request | A problem in one batch only, such as a late calendar change |
| Instructions or knowledge | The Project's standing instructions, examples and files | The same edit or error across many batches |
| Feature or model | The tool or the model tier | The task needs a capability the setup lacks, or speed or cost is wrong |
| A workflow step | Who checks what, and when; one task split in two | Errors are rare but costly, or one step soaks up the effort |
| No Claude for the step | The step goes back to people or another tool | Well-aimed, tested fixes still miss the criteria |
Start with the smallest change that reaches the pattern, and move down the table only when the evidence says so. The last row is a legitimate result, not a failure.
Length and tone go wrong in every batch, so they belong in the Project's standing instructions. First you save a copy of the current instructions. Then you replace the week-three "full details" line with the five criteria, plus the reason for the length limit: people read these posts while scrolling. You also add three approved posts as examples, which Anthropic's prompting guide calls one of the most reliable ways to steer format, tone and structure.
The unverified claims get two changes. An approved fact sheet of the claims legal has cleared goes into the Project knowledge, and the instructions say to use no others. Because this error is rare but costly, you also add a workflow step: Felix checks every claim about an attraction against the sheet before scheduling.
Engagement is an outcome, so you'll watch whether it follows the other fixes. And nothing points to the model: a more capable one would follow the week-three instruction just as faithfully. Model changes suit other feedback: Anthropic's prompt engineering guide notes that not every failing criterion is best fixed by prompting, and that speed or cost is sometimes easier to improve with a different model.
7.2.5 Test the change on the same inputs
The next batch looks much better. Is it? Its events are different, a quiet week flatters any setup, and Claude's drafts vary from run to run anyway. A better-looking batch is a hint, not a result.
The way to know is a small controlled comparison: run the same inputs through the old setup and the new one, and judge both against the same criteria. With everything else held still, any difference comes from your change. It's how an optician tests lenses: same chart, same distance, "clearer with one, or with two?"
You take four past calendars, including a busy holiday weekend and a week when a festival was cancelled, because awkward weeks show a setup's limits. A second Project holds the old instructions from your saved copy, and both Projects draft the same 32 event posts. Felix and Bronwen check all 64 posts against the five criteria, shuffled and unlabelled, so neither can tell which setup wrote which.
A controlled comparison
Old setup
New setup
The new setup also makes no claim outside the fact sheet. Because it bundles several changes, the result shows that the package works, not which part did it, and that is enough to decide whether to keep it. Anthropic's guide to success criteria and testing is written for developers building on its API. Its core advice still carries over: specific, measurable criteria, test cases that mirror the real work, and a comparison against the earlier version.
Outcome measures come later and need care, because engagement moves for many reasons, from the weather to how a platform ranks posts. So you watch it for six weeks and compare event posts with the spotlights and tips, whose length and tone you left alone. If event engagement recovers while the others stay flat, your change is the likeliest reason.
7.2.6 Close the loop, and know when to stop
The new setup goes live, and two jobs remain. First, record each change with the feedback behind it and the evidence for it, in a log kept outside the Project so Claude never mistakes notes about old rules for current ones. Look at the Why and Evidence lines: without them, the next person to meet the 45-word rule may undo it the first time someone asks for more detail.
Weekly posts Project: change log
Changed: event-post instructions rewritten to the five agreed criteria; three approved posts added as examples; approved fact sheet added; the week-three "full details" line removed.
Why: about 70% of Felix's edits over eight weeks were length and tone on event posts; legal flagged two unverified claims; criteria agreed by Imogen with Felix and Bronwen.
Evidence: four past calendars, old versus new setup, checked unlabelled against the criteria. Old: 9 of 32 event posts passed. New: 29 of 32.
New check: Felix confirms every claim about an attraction against the fact sheet before scheduling.
Watch: event engagement against the other posts for six weeks; Felix's editing time per batch (was about 40 minutes).
Second, close the loop with the people who gave the feedback. Tell Felix, Bronwen and legal what changed, what didn't and why, and what you'll watch. People who see their feedback acted on keep giving it. People who hear nothing stop, and your early warning goes quiet.
Sometimes the results point the other way. One weekly tip is an access post: step-free routes, accessible toilets, quiet hours at an attraction. Felix rewrote seven of the last eight, because access details change and each had to be checked with the venue. Adding each venue's approved access statement to the knowledge didn't help: the only safe post quoted the statement word for word, so Claude's drafting added nothing but risk.
That is a poor fit, and it usually shows one of three signs:
- Fixes have stopped helping. Well-aimed, tested changes at the right level still leave the results short of the criteria.
- Checking costs more than it saves. Reviewing and correcting take longer than doing the step without Claude.
- The step needs what Claude can't have. Information that isn't written down anywhere Claude can reach, or a judgement that must stay with a person.
Anthropic's AI Fluency course, on deciding which work to hand to AI, is clear that the goal isn't to automate everything. So the access posts leave Claude's batch; the team posts each venue's statement with a short introduction of their own. Claude keeps the other 19 posts, and the log records why.
7.2.7 The exam traps
Every trap here is a way of reacting to feedback without treating it as evidence.
- ✗ Rewriting the shared instructions after one complaint. ✓ Check it against the criteria and look across many outputs for a pattern first. One comment may be taste; a single serious error is the exception.
- ✗ Rewording the prompt again and again for reviewers who disagree. ✓ Take the conflict to the person accountable for the output and agree shared acceptance criteria. No wording meets two contradictory targets.
- ✗ Asking Claude to balance the stakeholders and pick a compromise. ✓ Let Claude sort the feedback and lay out the conflict; the accountable person decides what good means.
- ✗ Switching to a more capable model because reviewers keep editing. ✓ Change the level the pattern points to. Recurring edits usually call for instructions, examples or knowledge.
- ✗ Judging a change by next week's batch, or by Claude's opinion of it. ✓ Run the same inputs through the old and new setup and check both against the same criteria.
- ✗ Dropping Claude from the whole workflow because one step keeps failing. ✓ Take out the step the results show is a poor fit, keep what works, and record why.
7.2.8 Put it together: turn feedback into a tested change
You now have the whole cycle: read the pattern, have the owner settle the criteria and make the change where the pattern lives. Then prove it with a controlled comparison, record it and tell people. The quickest way to believe it is to watch a tested fix lose one ingredient.
Optimising workflows for efficiency (7.3) starts where this lesson ends: once feedback has brought a workflow up to its criteria, you can make it faster and cheaper. Your comparison set and change log are how you'll know an efficiency change hasn't quietly cost you quality.
Key takeaways
- ✓ Feedback is evidence: reviewer edits, stakeholder comments, error or rejection rates, time spent fixing and outcome measures each show part of the picture.
- ✓ Look for a pattern across many outputs before acting on one comment, but act on a single serious error.
- ✓ Conflicting feedback usually means the acceptance criteria were never agreed; the accountable person decides them, not Claude.
- ✓ Match the change to where the pattern lives: the prompt, instructions or knowledge, feature or model, a workflow step, or no Claude for that step.
- ✓ Test a change by running the same inputs through the old and new setup, judged against the same criteria, before you rely on it.
- ✓ Record what changed, why and on what evidence, and tell the people who gave the feedback.
- ✓ When tested fixes keep missing the criteria or checking costs more than it saves, stop using Claude for that step and keep the rest.
Check your understanding
4 questions written for this lesson, then one from the CCAO-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
36 CCAO-F questions on Domain 7, free
Every question in the bank is tagged to a domain, so you can drill 36 questions on Troubleshooting and Optimization alone, or sit the full 60-question timed simulator.
Open the CCAO-F question bank → Back to Domain 7 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.