Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCAR-P › Domain 1 › Lesson 1.5

CCAR-P · Domain 1 · 17% of the exam · Lesson 1.5 · 21 min read

Decomposition: cutting complex work into steps you can test

How to cut a complex task into steps Claude does well: where to cut, typed handoffs with provenance, gates, the human checkpoint and how many steps to use.

Written against objective 1.5 of the official CCAR-P exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

1.5.1 Why one prompt cannot review 400 contracts

Castellan Ware is a mid-size law firm acting for a grocery group that is buying Brightholme Foods, a regional supermarket chain. Before the client signs, it needs to know which of Brightholme's 400 supplier contracts let a supplier walk away, reprice or claim damages when Brightholme changes hands. Radomir, the partner running the deal, has three weeks and a stretched team.

The first attempt was the obvious one. An associate loaded a few dozen contracts at a time into one long request and asked Claude for "a risk register of change-of-control issues". The answers read well until a senior associate checked a sample. Two findings quoted a clause from a different contract, a termination right hidden in a schedule was missing, and nothing showed whether the quiet contracts had been read or skimmed. A second run produced a different register. The only thing anyone could test was the whole answer, and it could not be trusted.

This was not a failure of intelligence. One request did four jobs at once: sort the contracts, find the clauses, judge them against the client's rules, summarise. Each job got far more text than it needed, and the prose that came back had no point where anyone could check the work. Radomir asks Ysolde, the firm's solution architect, to turn the review into something a partner can sign. Her tool is decomposition: cutting a complex problem into smaller steps, each with its own focused prompt, a defined input and output, and its own test, joined together by code.

Per contract, her design has three steps. Classify it by type; extract four clause types (change of control, assignment, termination and liability); then check each clause against the client's risk playbook, the written rules for what the client will accept. A final step builds the risk register once a lawyer has confirmed every high-risk finding.

One request versus a decomposed review

One request

Dozens of contractsplus the whole playbook
Sort, find, judge, summariseall at once
One register in proseno way to check it

Decomposed review

Classifyone contract at a time
Extract, then checktyped output, tested per step
Lawyer confirms high riskthen code builds the register
The single request mixes four jobs over all the text and returns prose nobody can check; the decomposed review gives each job its own step, input and test.

1.5.2 What a step boundary buys, and what it costs

It is tempting to think decomposition works because small tasks are easier. On current models, that is the weaker half of the story. Anthropic's prompting guide notes that with adaptive thinking (Claude deciding how much to reason before it answers) and subagents (helper instances it can hand work to), Claude now handles most multistep reasoning internally. Explicit chaining, one API call per step, still earns its place when you need to inspect intermediate outputs or enforce a specific pipeline structure. So the real test of a cut is what the boundary lets you see or control.

A boundary buys four things. Focus: the extraction step sees one contract and four clause definitions, not a stack of contracts and a playbook; Anthropic's agent-building guide notes that models generally do better when each consideration gets its own call. A test point: a step with a defined input and output can be scored on its own. A choice of model or tool: sorting contracts by type suits a fast, low-cost tier such as Haiku, judging a liability cap deserves a stronger model, and counting findings needs no model at all. A place to check: between steps, code can inspect the output, retry it or hold it for a person.

Every boundary also has a price. Each call adds latency, and in a chain the waits add up. Each resends instructions and context, so cost grows with the step count. Errors compound: five chained steps that are each right 95% of the time, failing independently, are right together only about 77% of the time. And a careless cut can split what belongs together, such as a termination right on page nine and the definition of "change of control" on page two.

What a boundary buys and what it costs

Buys

Focusone job, only the context it needs
A test pointscore the step on its own
A choice of model or toolfast tier, strong tier or plain code
A place to checkretry, escalate or hold for a person

Costs

Latencywaits add up along a chain
Tokensinstructions resent on every call
Compounding errorfive steps at 95% is about 77%
Split contexta clause cut off from its definition
A cut is worth making only when what it buys outweighs its price for this workload.

1.5.3 Four ways to cut a problem

Knowing why to cut does not tell you where. There are four basic cuts, and real designs combine them.

  • By stage, a sequential chain: each step works on the previous step's output, as classify, extract and check do for one contract.
  • By independent part, which Anthropic calls sectioning: parts that do not depend on one another get a call each, and can run in parallel. Over many items of one kind, such as Brightholme's 400 contracts, this becomes map-reduce: the same step runs on every item (map) and a final step combines the results (reduce).
  • By expertise: a router sends each item to a prompt, model or tool built for its kind. Here the classify step routes logistics agreements, own-brand manufacturing deals and software licences to their own playbook sections.
  • Hierarchical: a sub-problem is itself decomposed, as the register splits into contracts and each contract into four clause questions.
Cut When it wins What it costs
By stage (chain) The work has a natural order and each stage needs the last one's result Latency adds up; an early error flows into every later step
By independent part (sectioning, map-reduce) Parts do not depend on each other, so they run in parallel with small contexts Links between parts are lost; the combining step needs its own design
By expertise (routing) Kinds of input differ enough that one prompt tuned for all serves each badly The router must itself be accurate; more prompts to test and maintain
Hierarchical The problem is large at more than one level More levels mean more handoffs to design and more places to lose the source

Memorise the four cuts and what picks each; distractors quietly ignore the costs column.

Nobody waits at a screen for Ysolde's map, so each stage runs as one batch through the Message Batches API, asynchronously and at half the standard price; most batches finish within an hour. Results come back in any order, so every request carries a custom_id naming its contract (and its clause, at the check stage). Code stores each step's output under that ID, so a failure reruns one step for one contract, not the whole review.

The reduce deserves the same thought. Counting, sorting and grouping findings by severity is deterministic, so code does it; the model writes only the one-page summary of top risks for Radomir. Asking a model to tally 400 contracts' findings would rebuild the overloaded request she has just taken apart.

1.5.4 Designing the handoffs between steps

Here is the failure that shows up a month into production: every step passes its own tests, and the register is still wrong. The fault usually sits between two steps. Suppose extraction returns a paragraph: "The agreement contains a change of control provision permitting termination". The check step must re-read prose, cannot tell an absent clause from one nobody looked for, and has no idea which page the sentence came from.

Treat each handoff as an interface contract, as you would between two services: each step receives only what its job needs and returns fields, not prose.

Step Receives Returns
Classify The contract's opening pages One type from a fixed list, or other
Extract One contract's full text and four clause definitions For each clause type: found, not_present or unclear, a verbatim quote and its location
Check One extracted clause and the playbook section for that contract type A risk level, the playbook rule applied and a one-line reason
Aggregate (code) Every checked finding, with the lawyer's decisions on flagged ones One register row per finding, with contract ID, quote, location, rule and reviewer

Two choices carry the weight. The explicit not_present status turns "we looked and it is not there" into an answer, so a clause type missing from the output means a failed step, never an absent clause. And provenance, the record of where each finding came from, is split by who knows it. Code attaches what code knows: the document ID, a hash of the document version, the prompt version and the model. The schema demands what only the model can supply: the exact quote and where it sits.

To make the shape binding, Ysolde uses structured outputs. She sets the request's output_config.format to type: "json_schema" with the schema below as its schema field, and the API constrains Claude's response to valid JSON that matches it. The lines that matter are the status enum and the required quote and location fields.

{
  "type": "object",
  "properties": {
    "clauses": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "clause_type": {"type": "string", "enum": ["change_of_control", "assignment", "termination", "liability"]},
          "status": {"type": "string", "enum": ["found", "not_present", "unclear"]},
          "quote": {"type": "string"},
          "location": {"type": "string"}
        },
        "required": ["clause_type", "status", "quote", "location"],
        "additionalProperties": false
      }
    }
  },
  "required": ["clauses"],
  "additionalProperties": false
}

1.5.5 Gates between steps, and where the lawyer belongs

Structured outputs guarantee the shape, not the truth: a well-formed quote can still be invented or copied from the wrong contract. So what stops a bad extraction flowing downstream? A gate: a check your code runs on a step's output before the next step may use it. Anthropic's agent-building guide describes this for prompt chains: programmatic checks that confirm the process is still on track. Think of the pass in a restaurant kitchen, the counter where the head chef looks at each plate before it goes out. In a pipeline, code does the routine looking, and a person stands at the pass only for the plates that matter.

Ysolde's pipeline has three kinds of gate:

  • EVERY CALL. The response's stop_reason field, which says why Claude stopped, must be end_turn, a normal finish. The docs warn that a refusal, or a reply cut off at max_tokens, may not match the schema.
  • AFTER EXTRACTION. Every found quote must appear word for word in that contract's text (after normalising whitespace), a string match that catches invented and misattributed quotes. All four clause types must be accounted for.
  • BEFORE AGGREGATION. All 400 contracts must have come back, since a request in a batch can end errored or expired.

An item that fails a gate is retried once, then sent to a stronger model, then queued for a person. It is never silently dropped.

Where the human checkpoint sits decides whether it works. Reviewing every step buries lawyers in 1,600 clause-level results and throws away the time saved. Reviewing only the final register is too late: a 400-row table is too big to check, and a wrong finding has already shaped the summary. The checkpoint belongs where judgment is consequential and before it spreads: after the playbook check, before aggregation, on every high-risk and unclear finding. Ysolde adds a small random sample of low-risk and not_present results, because a reviewer who only sees flagged items can never catch a missed one.

Gates, the checkpoint and the register

EXTRACTper contract, four clause types with quotes
GATEstop reason, quote in the text, four types
CHECKthe playbook section for the contract type
LAWYERhigh risk, unclear and a random sample
REGISTERcompleteness gate, then built in code
GATE → EXTRACT · fails a gate: retry, stronger model, then a person
Code gates every extraction and loops failures back to be retried or escalated; a lawyer sees high-risk and unclear findings, plus a sample, before code builds the register.

1.5.6 Settling the cut: an eval, and who draws the steps

If each step is easier, why not cut further, into one call per clause type or per page? Because every extra boundary pays its price whether or not it buys anything. Too few fail the other way: one overloaded prompt, no way to see where errors enter, and no cheaper model for the easy parts. Intuition cannot settle the count. An eval can: a set of real inputs with known right answers that you run each design against and score.

Ysolde's eval is 60 contracts a senior associate marked up by hand. Anthropic's evaluation guide asks for a set that mirrors the real task, edge cases included, so she adds awkward ones: clauses under odd headings, a side letter, a scanned schedule. She scores each design and each step on the metrics the engagement ranks. Recall of high-risk clauses (the share of real ones found) comes first, since a miss costs the client most, then false flags, cost per contract and time for the full run. She records the result for Radomir.

Decision: extraction and the playbook check run as separate steps; classification stays a separate step on a fast, low-cost tier.
Evidence (60 hand-labelled contracts): the combined extract-and-check call found 51 of 56 high-risk clauses; the split design found 55 of 56, at about 1.4 times the cost per contract. The full run still finishes overnight.
Rejected: one extraction call per clause type. Recall did not improve on the labelled set, and cost and latency rose.
Deciding requirement: a missed change-of-control clause is the costliest error in this engagement, so recall outranks cost.
Review: rerun the labelled set whenever the playbook, a prompt or a model changes.

The last question is who draws the steps. In static decomposition your code fixes the steps before any input arrives. In dynamic decomposition an orchestrating model reads each input and decides the subtasks itself, as in the orchestrator-workers pattern. Static wins when the steps are known in advance, as the four clause questions are: every contract takes the same path, so runs are predictable, cheaper and auditable. Dynamic paths vary from run to run, which makes them harder to test and to budget, so they pay off only when the subtasks cannot be predicted.

Some Brightholme contracts say "as amended by the side letter of 3 March", and which documents to fetch depends on the contract. So Ysolde keeps a static spine in code with one bounded dynamic step. When extraction returns unclear because of a cross-reference, a model-driven step gets read-only search over the data room (the deal's shared store of Brightholme's documents), a turn limit and the same output schema. Its result then passes the same gates as every other extraction.

1.5.7 The exam traps

Every trap here either skips decomposition, overdoes it, or leaves the space between steps undesigned.

  • ✗ Sending the whole job to one call, then reaching for a bigger model or longer context when it misses things. ✓ Decompose into focused, testable steps; a bigger model still returns one output with no place to check it.
  • ✗ Cutting as finely as possible because smaller steps are easier. ✓ Cut only where a boundary buys something, and let an eval on labelled inputs settle the count.
  • ✗ Passing free text between steps and leaving the source behind. ✓ Use typed handoffs with an explicit "not present" answer and provenance on every item.
  • ✗ Treating schema-valid output as correct. ✓ Gate in code between steps: stop reasons, verbatim quotes, completeness. The schema guarantees shape, not truth.
  • ✗ Reviewing everything, or only the final report. ✓ Place the human checkpoint on consequential findings before they spread, plus a sample of clean results to catch misses.
  • ✗ Letting an orchestrating agent plan every run when the steps never change. ✓ Draw known steps in code, and keep model-planned steps for what cannot be predicted, bounded by tools, turns and schema.

1.5.8 Put it together: decompose a review and measure the cut

You now have every piece: where to cut, what crosses each boundary, where the gates and the lawyer sit, and how an eval settles the count. Build a small version, race it against a single call, and watch one interface decision cause a failure.

Business value alignment (1.6) turns cost per contract and lawyer hours saved into the case for a design like this. Context window techniques (2.4) decide what each step's prompt carries. Evaluation datasets (4.2) turn a labelled set like Ysolde's into a lasting test suite, and human-in-the-loop validation (5.3) goes deeper on the reviewer's checkpoint.

Key takeaways

  • ✓ Decomposition cuts a complex problem into steps with focused context, a defined input and output, and a test each, joined together by code.
  • ✓ Cut where a boundary buys focus, a test point, a different model or tool, or a place to check, and weigh that against added latency, cost, compounding error and split context.
  • ✓ Cut by stage, by independent part (map-reduce over document sets), by expertise or hierarchically, and do deterministic aggregation in code.
  • ✓ Design each handoff as an interface: a schema for the output, an explicit "not present" answer and provenance on every item.
  • ✓ Gate every step in code, because a schema guarantees shape, not truth, and place the human checkpoint on consequential findings before they spread, with a sample of the rest.
  • ✓ Settle granularity with an eval on labelled inputs, draw known steps in code, and leave dynamic decomposition to a bounded orchestrating model only where subtasks cannot be predicted.

Check your understanding

4 questions written for this lesson, then one from the CCAR-P question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

33 CCAR-P questions on Domain 1, free

Every question in the bank is tagged to a domain, so you can drill 33 questions on Solution Design & Architecture alone, or sit the full 63-question timed simulator.

Open the CCAR-P question bank → Back to Domain 1 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources