Home › Study guides › CCDV-F › Domain 2 › Lesson 2.2
CCDV-F · Domain 2 · 33.1% of the exam · Lesson 2.2 · 22 min read
The life cycle of a Claude system, from pilot to retirement
How a Claude application lives from pilot to retirement: evals as tests, prompt and model changes as releases, staged rollout, monitoring and migrations.
Written against skill 2.2 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.
2.2.1 What happens after the demo works
Oksana is the only developer at Calder & Voss, a mid-sized accounting firm. Every month its clients send in thousands of supplier invoices, PDFs and phone photos alike, and staff type the supplier, number, date, amounts and tax into the bookkeeping system by hand. In two weeks Oksana builds a pilot: a small service that sends each invoice to Claude with a prompt and gets the fields back as JSON. On forty invoices from last month, every field comes back right. Bram, the partner who sponsored it, wants it live for all clients next month.
The demo is the easy part. The service will run for years, and three things will change under it. New suppliers and new layouts arrive. The prompt changes with every fix for an awkward invoice. And the model changes, because Anthropic retires older models as new ones arrive, and requests to a retired model fail. A script that passed forty invoices cannot prove that an edit made things better rather than worse. It cannot notice when it starts getting totals wrong, and it has no plan for the day its model disappears.
Software engineering met the general version of this problem long ago. The systems life cycle is the sequence every IT system goes through: plan it, develop it, deploy it (the implementation stage), operate it, maintain it and one day retire it. A Claude application goes through the same stages, with one difference that runs through all of them. Part of the program is now a prompt and a model. The model's output is not fixed, and the model has a retirement date that someone else sets.
How a team moves through the stages is set by its life-cycle framework. Waterfall walks the stages once, in order, and suits a system whose requirements barely change. Iterative frameworks, agile among them, go round the cycle in short loops and ship small changes often. DevOps brings development and operations together around an automated pipeline, so what operation learns flows straight into the next release. A Claude system belongs at the iterative end, because its inputs, its prompt and its model all keep changing.
The life cycle of a Claude system
2.2.2 Plan and develop: evaluations are your test suite
Here is the question that stalls many first Claude projects: how do you test a component whose output is not fixed? A normal function returns the same result for the same input, so a unit test can assert on it. Send Claude the same invoice twice and the replies can differ in small ways. Worse, a prompt edit that fixes one supplier's layout can break another's, and no failing build warns you.
The answer is an evaluation, or eval: a fixed set of realistic inputs with known correct outputs, run through the system and scored automatically. For the invoice service that means a golden set of 300 real invoices, with client details removed, whose correct fields an accountant has checked. Think of it as a standard exam paper with a fixed marking scheme. Every version of the service sits the same paper, so two scores can be compared, and a version that scores lower has visibly got worse. The eval plays the role a test suite plays in ordinary software: you run it before launch, and again on every change for the life of the system.
The plan stage decides what the eval measures. The targets come out of the requirements; what matters here is that they are specific and measurable, as Anthropic's guide to building evaluations recommends. "Extracts invoices well" cannot be tested. "At least 98% of fields exactly right on the golden set" can. Most systems need several criteria at once, because a release that is more accurate but twice as expensive can still fail. Cost is counted in tokens, the small pieces of text the model reads and writes; the API bills by them, and images are counted in tokens too.
| Criterion | Target for the invoice service | How it is graded |
|---|---|---|
| Field accuracy | 98% of fields exactly right | Code: exact match per field |
| Arithmetic | Net plus tax equals total on every output | Code: a sum check |
| Hard cases | 95% on credit notes, multi-page and foreign-currency invoices | Code, scored per group |
| Cost | At most 2 cents per invoice | Token counts from each response |
| Latency | Under 20 seconds per invoice | Timing in the eval run |
Memorise the shape, not the numbers: every criterion is a number a script can compute.
The golden set follows the three design principles in that guide. It mirrors the real mix of invoices, awkward ones included: credit notes, two tax rates on one invoice, a blurred photo, a document that is not an invoice at all. It is graded automatically wherever possible. And it favours volume, because many cases graded automatically beat a few graded by hand. The guide ranks the grading methods too. Code comes first, because an exact match on a field is fast and unambiguous. LLM-based grading comes next, for judgements code cannot make: a second model call scores the answer against a clear rubric. Human grading comes last, because it is slow and expensive.
One more rule comes from ordinary testing: keep the golden set out of the prompt. An invoice used as a prompt example no longer tells you how the service handles invoices it has not seen.
2.2.3 Deploy: every prompt or model change is a release
Three months after launch, a client's biggest supplier starts printing two tax lines on its invoices, and the service reads only the first. Bram asks Oksana to "just tweak the prompt". It is tempting to open the production prompt and edit it on the spot: no code change, no build, no deploy. Resist it. In a Claude system the prompt IS part of the program, and so are the model ID (the exact model name your code requests), the request settings, the tool definitions and the output schema. A one-line prompt edit can change how every invoice is read.
So each of those changes is a release: a versioned bundle that ships as a unit and takes the same path as a code change. How you store and version the bundle is a configuration question. What matters here is the path, and its first step is the eval gate: score the golden set on the live release and on the candidate, and block the candidate if it scores lower. In this sketch, look at the two lines that take the model and the prompt from the release, and at the last three lines, which are the gate.
def field_accuracy(release, cases):
correct = 0
for case in cases:
reply = client.messages.create(
model=release["model"], # the model ID is part of the release
system=release["prompt"], # and so is the prompt
max_tokens=4096,
messages=[{"role": "user", "content": case["invoice"]}],
)
text = next((b.text for b in reply.content if b.type == "text"), "")
fields = parse_json_or_empty(text) # bad JSON scores zero, not a crash
correct += sum(fields.get(f) == case["expected"][f] for f in FIELDS)
return correct / (len(cases) * len(FIELDS))
baseline = field_accuracy(LIVE_RELEASE, golden_set)
candidate = field_accuracy(NEW_RELEASE, golden_set)
if candidate < baseline: # the gate: no release may score lower
sys.exit(f"Blocked: {candidate:.1%} against {baseline:.1%} for the live release")
A real gate would also allow a small margin, because scores vary a little from run to run. It would set a floor for each edge-case group too, so a good average cannot hide a collapse on credit notes. Oksana adds 25 two-tax-line invoices to the golden set, edits the prompt, and the gate passes. Then the release goes out in stages, not all at once.
The path of every release
In shadow mode the new release reads real invoices alongside the live one, and your code compares the two outputs without using the new one. It doubles the API cost of the traffic you shadow, but it tests on today's invoices rather than last year's golden set. In a canary, the new release handles a small share of real traffic, say 5%, watched by the same signals as production. The name comes from the canaries miners once carried underground: a small, early warning before the danger reaches everyone.
Ramping up gradually also protects the first launch: Anthropic warns that a sharp jump in usage can trip acceleration limits and return 429 (rate limit) errors.
Rollback means switching production back to the last release that worked, and it is what makes all of this safe. Keep the previous release deployable as a whole bundle, so the old prompt comes back with the parser that expects its output. Switch releases by changing configuration, not by rewriting code. And write down the rollback criteria before the rollout starts: for example, a staff correction rate above 3% or a cost per invoice above budget. Deciding in advance turns rollback into a routine step instead of an argument in the middle of an incident.
2.2.4 Operate: a service can be up and still wrong
Once the service is live, how do you know it is healthy? For ordinary software the usual dashboards answer that: uptime, response time, error rate. For a Claude service they can all be green while it writes the wrong total into a client's books, because a valid JSON reply says the call worked, not that the invoice number is right. It is like judging a courier only by whether the van arrives on time: uptime measures the van, not the parcel. Operating a Claude system means watching two more things next to the usual ones: quality and cost.
| Signal | What it catches | Where it comes from |
|---|---|---|
| Validation failures | Malformed JSON, missing fields, totals that do not add up | Your checks on every response |
| Correction rate | Fields staff had to fix before posting | The bookkeeping review screen |
| Sampled grading | Slow drift nobody reports | A weekly sample scored like the eval |
| Tokens and cost per invoice | A growing prompt, larger inputs, a pricier model | usage on each response; the Usage and Cost API |
| Latency and errors | Capacity problems, traffic spikes, 429 (rate limit) and 529 (overloaded) errors | Your request logs |
The first three rows are quality signals, and they exist only if you build them. Put an alert on each row, and log with every result the release that produced it, so a rise in corrections can be traced to a prompt version or a model.
A year of operation at Calder & Voss shows why. At quarter end the volume tripled, but cost per invoice stayed flat, the figure that tells growth from waste. In the summer one supplier redesigned its invoices, and its correction rate rose from 1% to 6%; the weekly sample flagged it before any client complained. In the autumn, cost per invoice crept up 40%: a large client had started sending twenty-page month-end statements instead of single invoices, and the cost alert caught it before the bill did. Each failure became a new eval case, and each fix went out along the release path. That is how maintenance feeds back into development.
For cost, the usage object on each response gives its input and output token counts, so your code can log the cost of every invoice. The Usage and Cost API reports the whole organisation's usage by model, API key or workspace. A spend limit in the Claude Console caps the monthly bill, but requests fail once it is reached, so it is a backstop, not a monitor. And quality monitoring never ends at launch: even with the model ID unchanged, Anthropic notes that updates to its serving infrastructure can occasionally cause minor differences in behaviour.
2.2.5 Maintain and retire: when the model reaches end of life
Fourteen months after launch, an email arrives from Anthropic: the model the service runs on is deprecated, and a retirement date is set a few months out. This maintenance event catches teams out, because nothing in their own code changed. Every Claude model moves through the same stages, and Anthropic's model deprecations page names them.
| Status | What it means | What your team does |
|---|---|---|
| Active | Fully supported and recommended | Build on it; note its tentative retirement date |
| Legacy | No longer receives updates; may be deprecated later | Start testing a newer model |
| Deprecated | Still works, not recommended; has a replacement and a retirement date | Migrate before the date |
| Retired | No longer available | Too late: every request to it fails |
Know the four names and the rule that matters most: a deprecated model still answers, a retired one does not.
Anthropic gives at least 60 days' notice before retiring a publicly released model, by email and in the documentation, and warns that deprecated models are likely to be less reliable than active ones. Until then the model will not shift under you: a model ID names a fixed snapshot, and Anthropic ships any update under a new ID. So the migration is a release you make, on your own schedule, inside the notice window. It is like an expiry notice for a bank card: the old card works until the date, and the real work is finding every subscription still charging it before the payments start to bounce.
Treat the migration like any other release, with two extra steps at the front. First, find every caller: the Usage page in the Claude Console exports usage by API key and model, and at Calder & Voss it turned up a forgotten nightly script still calling the old model. Second, read the replacement's migration guide, because a migration is rarely a one-line ID swap. The guide for Claude Sonnet 5.5, for example, lists five request settings that some earlier models accepted and that it rejects with a 400 error. Among them are a non-default temperature (a setting that makes the wording more or less random) and a tool_choice that forces a tool call.
Oksana's replacement model had a change of exactly this kind. The service had sent a temperature of 0 since the pilot, and the replacement rejected it, so the first eval run failed on every invoice. That discovery cost a few cents; on retirement day it would have stopped the firm's bookkeeping. The cost baseline moved too. The same guide warns that, compared with Claude Sonnet 4.6 or Claude Haiku 4.5, the same text produces about 30% more tokens. A 2000 by 1500 pixel scan costs about 2.5 times as many tokens. The golden set's cost per invoice jumped accordingly, and Oksana agreed a new budget with Bram before the canary began.
Then the release went through shadow, canary and full traffic while the old model still existed as a rollback target. That is the real reason to start early: after the retirement date there is nothing to roll back to.
The system itself will retire one day too, and that stage has its own checklist, not just a switch. Move its traffic to the successor and revoke its API keys. Delete stored invoices and logs under the firm's retention rules. Keep the golden set and the release history: they are the most valuable things a successor inherits.
2.2.6 The exam traps
Almost every trap here skips a stage of the cycle: it ships without the gate, watches without quality signals, or waits until the model is gone.
- ✗ Editing the production prompt directly because no code changed. ✓ Treat the prompt as part of the release: eval gate, staged rollout. A one-line edit can change every output.
- ✗ Approving a change after trying a handful of inputs by hand. ✓ Score the full eval set, edge-case groups included, against the live release. A few inputs tell you only about those inputs.
- ✗ Calling the service healthy because uptime, latency and errors look fine. ✓ Also monitor quality and cost per unit of work. A Claude service can succeed on every call and still be wrong.
- ✗ Sending all traffic to a new prompt or model at once. ✓ Shadow, then canary, then full, with rollback criteria agreed first and the previous release kept deployable.
- ✗ Treating a model migration as an ID swap made near the retirement date. ✓ Start when the notice arrives: find every caller, check breaking changes, rerun evals and cost, and keep the old model as a rollback while it exists.
Four shortcuts, one release path
2.2.7 Put it together: ship a change, then roll it back
You now have the whole cycle. Success criteria and an eval set come before launch; releases pass a gate and go out in stages; operation watches quality and cost; migrations start the day the notice arrives. The quickest way to make it stick is to build a tiny release gate and watch it work.
Several skills build on this cycle. Configuration management (2.6) is how you pin the model and version the prompt, so a release is one reproducible bundle you can roll back to. Debugging and error handling (4.1) takes over when monitoring raises an alarm and you must find whether the fault lies in your integration or in the model's output. Model selection and tradeoffs (5.3) looks closely at the behaviour changes between model releases that every migration tests for. Cost and token management (5.4) turns the cost signal into a budget you can model.
Key takeaways
- ✓ A Claude system goes through the usual life cycle (plan, develop, deploy, operate, maintain, retire), best run as an iterative loop, because its behaviour lives in the prompt and the model as well as the code.
- ✓ Evals are the test suite: realistic inputs with known answers, scored automatically against measurable success criteria set before launch.
- ✓ Every change to the prompt, model ID, settings or schema is a release: it passes an eval gate, goes out through shadow and canary stages, and can be rolled back by configuration.
- ✓ In operation, monitor quality and cost per unit of work as well as uptime, latency and errors, and turn every production failure into an eval case.
- ✓ Models move from active to legacy, deprecated and retired; Anthropic gives at least 60 days' notice for publicly released models, and requests to a retired model fail.
- ✓ A migration is a release, not an ID swap: find every caller, check breaking changes, rerun evals and the cost baseline, and roll out while the old model still exists as a rollback.
Check your understanding
4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.
51 CCDV-F questions on Domain 2, free
Every question in the bank is tagged to a domain, so you can drill 51 questions on Applications and Integration alone, or sit the full 53-question timed simulator.
Open the CCDV-F question bank → Back to Domain 2 →
The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.