Claude Certification Program · v1.0 · Effective July 2026 · All four tracks open

Home › Study guides › CCDV-F › Domain 5 › Lesson 5.2

CCDV-F · Domain 5 · 16.8% of the exam · Lesson 5.2 · 21 min read

SDKs, REST and websockets: the plumbing around every Claude call

What the Claude SDKs do on top of the REST API, when to call it directly, SSE versus websockets for the browser, and timeouts, retries and request ids.

Written against skill 5.2 of the official CCDV-F exam guide (Version 1.0, effective July 2026). An independent resource, not affiliated with Anthropic; the practice questions are written from scratch.

5.2.1 Why a working API call is not yet a working feature

Anneke builds the Python back-end of a small startup's newsletter editor, and Farid builds its React front-end. Their new feature is a writing assistant. A writer highlights a paragraph, clicks "Make it friendlier" or "Draft an intro", and a suggestion appears in a sidebar. Anneke's first version was a short Python script. It sent the paragraph to Claude with a plain HTTP library, waited for the reply and returned it. It did not use Anthropic's SDK (software development kit), the official library for calling the API. On her laptop it worked every time.

Inside the editor, the script met real writers. They stared at an empty sidebar for fifteen seconds, then the whole suggestion landed at once. One busy lunchtime, an overloaded error from the API became a blank "Something went wrong", and writers who double-clicked got two different suggestions stacked in one sidebar. Once Anneke added streaming (sending the suggestion piece by piece as Claude writes it), a writer on a train watched half a paragraph arrive and then nothing, forever. And when one odd reply needed investigating, the logs held nothing Anthropic support could look up.

None of these is a model problem. They are plumbing problems, and they live on two hops. Between Anneke's server and Claude, the transport is fixed. The Claude API is a REST API, plain HTTPS requests answered in JSON, and it streams a reply as server-sent events: small messages written one after another into a single HTTP response. Between her server and Farid's editor, the transport is the team's own choice. Around every call on either hop sit the same engineering habits. Decide how long to wait and when to try again, make a repeated click harmless, stop a fast sender from swamping a slow reader, and log what lets you trace a failure.

Two hops between a click and a suggestion

React editorthe writer clicks "Draft an intro"
Websocketthe team's choice of transport
Python serverthe SDK and the API key live here
HTTPS and server-sent eventsfixed by the Claude API
Claude APIPOST /v1/messages
Claude's side of the path is fixed, HTTPS and server-sent events through the SDK; the browser's side, and every habit around each call, is yours to design.

5.2.2 The REST call underneath every SDK method

Here is the question that puzzles developers new to large language models: if the API is only HTTPS and JSON, why install a library at all? To answer it, look at what the library hides. The Claude API is a RESTful API at https://api.anthropic.com, and a Messages call is a POST to /v1/messages with a JSON body. Three headers travel with it: your API key as Authorization: Bearer <key> (the older x-api-key header still works), anthropic-version with a value such as 2023-06-01, and content-type: application/json. The reply is JSON holding the content blocks, stop_reason and usage. Every response also carries a request-id header, the unique id Anthropic support asks for when you report a problem with a request.

Anneke's first version was exactly this call, written with httpx, a popular Python HTTP library. Look at the headers she sets by hand, the timeout she has to choose, and the last three lines, where a failure is only an exception and the reply only a dictionary.

import os
import httpx

resp = httpx.post(
    "https://api.anthropic.com/v1/messages",
    headers={
        "authorization": f"Bearer {os.environ['ANTHROPIC_API_KEY']}",  # AUTH: yours to send
        "anthropic-version": "2023-06-01",             # REQUIRED on every request
        "content-type": "application/json",
    },
    json={"model": MODEL, "max_tokens": 1024,
          "messages": [{"role": "user", "content": prompt}]},
    timeout=60.0,                                      # httpx's own default is 5 seconds
)
resp.raise_for_status()                    # a 529 is just an exception now: no retry
data = resp.json()                         # a plain dict: no types to catch a typo
log.info("request-id=%s", resp.headers["request-id"])

Everything this call does not do is now her job. A 529 (the API's "overloaded" status) at lunchtime reaches the writer as an error, and a typo such as data["stop_reson"] fails only when that line runs. With "stream": true she would also parse the events herself. That means named events in a fixed order, ping events to skip, new event types that must not crash her parser, and an error event that can follow a 200 status. Anthropic's streaming guide recommends the SDKs for streaming and warns that a direct integration must handle these events itself.

Calling REST directly still has its place. Official SDKs exist for Python, TypeScript, C#, Go, Java, PHP and Ruby, so a service written in another language talks plain HTTP. A curl command that prints the response headers is the quickest way to reproduce a bug with a request-id you can hand to support. Think of it as doing your own tax return instead of using tax software: legal and sometimes the right call, but every rule the software would have applied is now yours to remember.

5.2.3 What the SDK adds, and which defaults to change

It is tempting to treat the SDK as a thin convenience whose defaults suit every call. Both halves of that belief cause trouble: the SDK does real engineering work, and its defaults have to suit every kind of use, not your editor sidebar.

What the SDK does Its default What you still decide
Sends the auth, anthropic-version and content-type headers Reads ANTHROPIC_API_KEY from the environment Keeping the key on the server
Retries connection errors, 408, 409, 429 and 5xx with exponential backoff, honouring retry-after 2 retries (max_retries) Whether another layer also retries
Times out requests 10 minutes, and a timeout is retried too A timeout per call that matches what the user will wait for
Types requests and responses, and raises a typed exception per status RateLimitError, APIStatusError, APITimeoutError and more Which errors to show, retry or fall back on
Parses the event stream and assembles the final message messages.stream(), text_stream, get_final_message() What happens when a stream fails halfway
Exposes each response's request-id _request_id on response objects Logging it next to your own ids

Memorise what the SDK takes off your hands and its two headline defaults, two retries and a ten-minute timeout. Recognise the class names when you see them.

Here is the arithmetic that surprises teams. The default timeout is ten minutes, and a request that times out is retried twice. A call stuck on a bad network can therefore hold a writer's click for about half an hour. The opposite mistake hurts too. A non-streaming call's connection sits idle until the whole reply is ready, so a timeout shorter than its normal duration fails every time, and each retry waits the full timeout again. Set timeouts from measured durations, and stream anything long.

Anneke's server gives the client a 60-second ceiling and gives short calls a tighter budget. Look at the with_options line, the typed exception with its request id, and the _request_id logged on success.

import anthropic

client = anthropic.AsyncAnthropic(  # reads ANTHROPIC_API_KEY, sends every required header
    timeout=60.0,                   # default is 10 minutes, and timeouts are retried
    max_retries=2,                  # the default: backoff on 408, 409, 429, 5xx, network errors
)

async def suggest_title(text: str) -> str:
    quick = client.with_options(timeout=15.0, max_retries=1)  # a SHORT call, a short budget
    try:
        msg = await quick.messages.create(
            model=MODEL, max_tokens=60,
            messages=[{"role": "user", "content": f"Suggest one title for:\n{text}"}],
        )
    except anthropic.APIStatusError as err:          # typed: a 4xx or 5xx, after any retries
        log.warning("title status=%s request_id=%s",
                    err.status_code, err.response.headers.get("request-id"))
        raise
    log.info("title request_id=%s", msg._request_id)  # the id Anthropic support asks for
    return msg.content[0].text

One gap remains. A timeout or a dropped connection produces no response, so there is no request-id to log; your own id for the click, logged before the call, is the only thread back to the writer. And if a job queue or a wrapper of your own also retries, the attempts multiply, so choose one layer to own retries.

5.2.4 Server-sent events in, websocket out

Here is the belief behind many over-built streaming features: streaming needs a websocket. It does not. Claude streams with server-sent events (SSE): your server makes one ordinary HTTPS request, and the API keeps the response open and writes small named events into it as the reply forms. SSE runs one way, from server to client, over plain HTTP. Your server never talks back to Claude mid-reply, so nothing more is needed on that hop, and the SDK's messages.stream() hides the events behind text_stream.

The second hop is where you choose. Your server can stream to the browser in the same style, as an SSE response from its own endpoint. Or it can use a websocket: a connection that starts as an HTTP request, is upgraded, and then carries messages both ways for as long as both sides keep it open. Think of SSE as a radio broadcast and a websocket as a phone call. A radio is perfect when one side talks and the other listens; a phone call earns its extra setup when both sides need to speak whenever they like.

Two ways to push tokens to the browser

SSE from your endpoint

One wayserver to browser only
An ordinary HTTP responsetext/event-stream
Stop means closing the request

fits one question, one streamed answer

Websocket

Both wayseither side sends at any time
One long-lived connectionmany drafts share it
Handshake auth, heartbeats, reconnectsyour duties

fits edits mid-stream and parallel drafts

SSE is the simpler fit when text flows one way; a websocket earns its extra duties when the browser must talk back while the text is still arriving.

For one question and one streamed answer, SSE is the lighter choice. The browser's built-in EventSource reconnects on its own but only sends GET requests, so pages that post a prompt often read the stream with fetch instead. The writing assistant needs more. A writer can ask for "shorter" while text is still arriving, stop one draft while another keeps going, and run a title suggestion beside an intro. The editor already keeps a websocket open for autosave, so assistant traffic shares it, with a draft id on every message so several drafts can use one connection.

A websocket brings duties SSE does not. Browsers cannot add custom headers to the handshake, so the server checks the session cookie and the Origin header, or a short-lived token, before accepting. Heartbeats stop proxies and load balancers from closing an idle socket, every proxy on the path must allow the upgrade, and the browser must reconnect and resend unfinished requests. One thing never changes: the SDK and the API key stay on the server. The TypeScript SDK refuses to run in a browser unless you set dangerouslyAllowBrowser, because a key that reaches the browser belongs to anyone who opens the developer tools.

5.2.5 The relay: backpressure, cancellation and repeat requests

Now the relay itself: the code on Anneke's server that reads Claude's stream and writes to the websocket. Three production failures live here, and each has a name.

Backpressure means letting a slow reader slow down a fast writer. Claude can produce text faster than a phone on a train can receive it, and a relay that reads everything at once piles the gap up in memory, one growing buffer per slow connection. Picture a supermarket checkout belt: when the bagger falls behind, the belt stops and the cashier waits. The simplest form is to await each send before reading the next piece; any queue between reader and sender gets a maximum size, so a full queue makes the reader wait. In the relay, look at the awaited send, and at the draft id logged before the stream opens: a stream that never opens has no request-id at all.

async def relay(ws: WebSocket, draft: str, prompt: str) -> None:
    log.info("draft=%s start", draft)           # YOUR id, logged before any response exists
    async with client.messages.stream(         # the stream lives inside this block
        model=MODEL, max_tokens=1500,
        messages=[{"role": "user", "content": prompt}],
    ) as stream:
        async for text in stream.text_stream:
            await ws.send_json({"draft": draft, "text": text})  # AWAIT: a slow socket slows the reader
        final = await stream.get_final_message()
    await ws.send_json({"draft": draft, "done": final.stop_reason})

Cancellation covers the writer who clicks Stop or closes the tab. The server runs each draft as its own task. Cancelling the task breaks out of the loop and exits the async with block, so the server stops reading at once. The TypeScript SDK's docs describe the same move as the way to cancel a stream: break from the loop, or call stream.controller.abort(). Without cancellation, the server keeps reading a stream nobody will see.

Repeat requests come from double clicks and reconnects. The Messages API is stateless, so a repeat corrupts nothing on Anthropic's side, but it is a new request: new tokens on the bill and a different text in the same sidebar. So each click gets a draft id, and the server starts a stream only for an id it has not seen. Sending the same draft twice now has the same effect as sending it once, which is what idempotent means, and the draft id is the idempotency key. Here the record is a dictionary per connection; in production it lives in a shared store, so a reconnect that lands on another server is caught too.

In the endpoint, look at the cancel(), the draft-id check and the finally that runs when the tab closes.

@app.websocket("/ws/assist")
async def assist(ws: WebSocket) -> None:
    await ws.accept()                          # only after checking the session cookie
    tasks: dict[str, asyncio.Task] = {}
    try:
        while True:
            msg = await ws.receive_json()      # raises when the tab closes
            if msg["type"] == "stop" and msg["draft"] in tasks:
                tasks[msg["draft"]].cancel()   # CANCEL: the relay stops reading at once
            elif msg["type"] == "draft" and msg["draft"] not in tasks:  # IDEMPOTENT on the id
                tasks[msg["draft"]] = asyncio.create_task(relay(ws, msg["draft"], msg["prompt"]))
    finally:
        for task in tasks.values():
            task.cancel()                      # tab closed: cancel every relay

One failure no setting prevents: an error event, such as overloaded_error, can arrive after the stream has started. Anthropic's error documentation says an error that arrives after the 200 status does not follow the usual handling, so do not count on the automatic retries to cover it. Here, a try around the block (not shown) sends a failed message, Farid's editor greys out the partial text, and Retry starts a new draft id. You may also resume from the partial text, as the streaming docs describe, but never replay the request and append a second answer to the first.

5.2.6 The exam traps

Every trap here is a demo habit carried into production. A question usually describes the symptom (a spinner, a frozen stream, a duplicate or a leaked key) and asks for the fix.

  • ✗ Calling Claude from the browser so tokens stream straight to the page. ✓ Keep the SDK and the key on your server and relay the stream. A key that ships to the browser can be read by anyone.
  • ✗ Assuming streaming requires a websocket. ✓ Claude streams with server-sent events. Choose a websocket for the browser hop only when the browser must send messages mid-stream.
  • ✗ Leaving the SDK's defaults on an interactive call. ✓ Set a timeout per call from measured durations. Timeouts are retried, so ten minutes can become half an hour.
  • ✗ Wrapping the SDK in your own retry loop as well. ✓ Let one layer own retries. Stacked layers multiply the attempts and the waiting.
  • ✗ Buffering the stream without limit and ignoring Stop and closed tabs. ✓ Await each send or bound the queue, and cancel the relay when the writer stops or leaves.
  • ✗ Replaying a stream that failed halfway and appending the new text. ✓ Handle a mid-stream error in your own code: discard the partial text or resume from it explicitly, and log the ids so the failure can be traced.

5.2.7 Put it together: build the relay and break it

You now have every piece. The REST call sits underneath, the SDK adds retries, timeouts, types and stream parsing, and the relay carries SSE from Claude to a websocket with backpressure, cancellation and idempotency. The quickest way to make it stick is to build the path, then break two of its habits and watch what happens.

This plumbing sits under much of what follows. Model selection and tradeoffs (5.3) decides which model sits behind this call, and the latency you can now measure is one of its inputs. Cost and token management (5.4) turns the usage and request ids you now log into budgets and alerts. Identity, secrets and key management (7.4) takes the rule that the key stays on the server much further.

Key takeaways

  • ✓ The Claude API is REST: a POST /v1/messages with a JSON body and the auth, anthropic-version and content-type headers, and every response carries a request-id.
  • ✓ Call REST directly only when no SDK fits, knowing you then own headers, retries, timeouts, stream parsing and error handling.
  • ✓ The official SDKs retry transient failures twice with backoff and time out after ten minutes; set both per call, because timeouts are retried too, and let one layer own retries.
  • ✓ Log your own id before every call and the request-id whenever a response arrives, because a timeout leaves only your own id.
  • ✓ Claude streams to your server as one-way server-sent events; for the browser hop, choose SSE for one-way text and a websocket when the browser must talk mid-stream, and keep the key on the server either way.
  • ✓ A relay needs backpressure (await each send or bound the queue) and cancellation when the user stops or leaves.
  • ✓ A client-generated id makes repeated requests harmless, and a stream that fails midway is discarded or resumed explicitly, never replayed and appended.

Check your understanding

4 questions written for this lesson, then one from the CCDV-F question bank on the same topic. Every answer option is explained, including the ones you did not pick. Nothing is stored.

27 CCDV-F questions on Domain 5, free

Every question in the bank is tagged to a domain, so you can drill 27 questions on Model Selection and Optimization alone, or sit the full 53-question timed simulator.

Open the CCDV-F question bank → Back to Domain 5 →

The question bank is free. It asks for an account only because the quiz engine has to store answers to score them and show which domains are weak. The questions on this page need nothing.

Sources