Home / Learn / Anthropic Python SDK: install, set base_url, stream, and handle errors

Anthropic Python SDK: install, set base_url, stream, and handle errors

The anthropic package is Anthropic's official Python client for the Claude API. It handles authentication, retries, streaming and typed errors, and it works unchanged against any Anthropic-compatible endpoint: set base_url (or ANTHROPIC_BASE_URL) and the rest of your code stays the same.

Install

terminal
pip install anthropic

The 1.x releases run on httpx2 and need Python 3.10 or newer. If you pass your own HTTP client, use the SDK's DefaultHttpxClient so its default timeouts and connection limits are kept.

Create a client

app.py
import anthropic

client = anthropic.Anthropic(
    api_key="your_key",   # default: ANTHROPIC_API_KEY
    base_url="https://aiprimetech.io",    # default: ANTHROPIC_BASE_URL, else https://api.anthropic.com
)

base_url is the host root. The SDK appends /v1/messages itself, so https://aiprimetech.io/v1 would produce /v1/v1/messages and a 404 page not found. With both variables set in the environment, anthropic.Anthropic() needs no arguments at all:

terminal
export ANTHROPIC_API_KEY=your_key
export ANTHROPIC_BASE_URL=https://aiprimetech.io

Send a message

app.py
message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Summarise TCP slow start in two sentences."}],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

print(message.usage.input_tokens, message.usage.output_tokens)
print(message._request_id)   # log this with any failure report

content is a list of blocks (text, thinking, tool use), so check block.type instead of reading content[0].text blindly.

Stream responses

stream.py
with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=16000,
    messages=[{"role": "user", "content": "Write a short story about a lighthouse."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()

Stream anything with long input, long output or a large max_tokens: a single non-streaming request has to finish inside the timeout, a stream does not. messages.create(..., stream=True) gives you the raw event iterator if you don't need the helper.

Async client

async_app.py
import asyncio, anthropic

client = anthropic.AsyncAnthropic(base_url="https://aiprimetech.io")
limit = asyncio.Semaphore(3)   # stay within your key's concurrency

async def ask(question: str) -> str:
    async with limit:
        msg = await client.messages.create(
            model="claude-haiku-4-5",
            max_tokens=512,
            messages=[{"role": "user", "content": question}],
        )
        return next(b.text for b in msg.content if b.type == "text")

async def main():
    answers = await asyncio.gather(*(ask(q) for q in ["What is DNS?", "What is BGP?"]))
    print(answers)

asyncio.run(main())

The semaphore matters. Firing hundreds of requests at once with asyncio.gather is the most common cause of Concurrency limit exceeded errors we see from Python clients.

Timeouts and retries

Defaults: a 10-minute timeout per request and 2 automatic retries with exponential backoff. The SDK retries connection errors, 408, 409, 429 and 5xx responses. Because timeouts are retried too, the worst-case wall-clock time is the timeout times the number of attempts.

app.py
client = anthropic.Anthropic(base_url="https://aiprimetech.io", timeout=60.0, max_retries=4)

# override for one call only
client.with_options(timeout=300.0).messages.create(
    model="claude-opus-5-5",
    max_tokens=4000,
    messages=[{"role": "user", "content": "Review this design doc..."}],
)

Error handling

Status Exception Retry?
400anthropic.BadRequestErrorNo: fix the request
401anthropic.AuthenticationErrorNo: wrong or missing key
403anthropic.PermissionDeniedErrorNo
404anthropic.NotFoundErrorNo: model ID or URL
429anthropic.RateLimitErrorYes, after a pause
5xx, including 529 overloadedanthropic.InternalServerErrorYes
Network failureanthropic.APIConnectionErrorYes
Timeoutanthropic.APITimeoutErrorYes, or stream
errors.py
try:
    message = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
                                      messages=[{"role": "user", "content": "Hi"}])
except anthropic.NotFoundError as e:
    print("Bad model ID or URL:", e.message)
except anthropic.RateLimitError as e:
    print("Rate limited, retry after", e.response.headers.get("retry-after"))
except anthropic.APIStatusError as e:
    print(e.status_code, e.type, e.message)
except anthropic.APIConnectionError:
    print("Network problem")

Catch the specific classes first and APIStatusError last; one broad except hides the difference between errors worth retrying and errors that need a code change. e.type gives the API's error type string, such as rate_limit_error or overloaded_error.

Model IDs

Model ID Use it for
claude-sonnet-5Default for coding and agent work: fast, strong, mid-priced.
claude-opus-5-5Harder reasoning, large refactors, long agent runs.
claude-fable-5-1Most capable and most expensive; save it for the hardest tasks.
claude-haiku-4-5Cheap and quick: titles, summaries, classification, small steps.

client.models.list() returns the models an endpoint serves, so you can check an ID before you use it. Use the IDs exactly as written: no date suffixes.

Troubleshooting on a gateway

What you see Cause Fix
NotFoundError with 404 page not found/v1 in base_urlUse the host root
AuthenticationErrorKey not set in the process, or pasted with whitespacePrint len(os.environ["ANTHROPIC_API_KEY"])
RateLimitError: Concurrency limit exceeded for userToo many requests in flightAn asyncio.Semaphore as above
APITimeoutError on long jobsNon-streaming request running past the timeoutStream, or raise timeout
InternalServerError: No available accountsTemporary upstream capacityLet the SDK retry; raise max_retries
Insufficient account balanceBalance emptyTop up

More on model-side errors: the Claude API errors guide.

Frequently asked questions

How do I install the Anthropic Python SDK?
pip install anthropic. Version 1.x needs Python 3.10 or newer.
How do I set a custom base URL in the Anthropic Python SDK?
Pass base_url to anthropic.Anthropic() or set ANTHROPIC_BASE_URL. Use the host root; the SDK appends /v1/messages.
What are the default timeout and retries?
10 minutes per request and 2 retries with exponential backoff, covering connection errors, 408, 409, 429 and 5xx.
How do I stream with the Python SDK?
Use with client.messages.stream(...) as stream: and iterate stream.text_stream; stream.get_final_message() returns the complete message.
Start using Claude in minutes

Get an API key — no Anthropic account or waitlist required.

Get your API key

AI Prime Tech is an independent API gateway. It is not affiliated with, endorsed by, or a reseller of Anthropic. Claude and related model names are trademarks of their respective owners.

More Claude API guides

How to Get a Claude API KeyClaude API Pricing, ExplainedThe Cheapest Way to Use the Claude APIUnlimited Cursor with the Claude APIUnlimited Claude Code AccessClaude API Gateway vs. Anthropic DirectUsing Claude Through an OpenAI-Compatible APIClaude models explained: a Claude model comparison of Claude Opus vs Sonnet vs Haiku vs Fable — all Claude API models, Anthropic model names and prices (2026)Paying for the Claude API with CryptoClaude API Access Without a WaitlistHow to Reduce Claude Token UsageUnlimited Cline with the Claude APIUnlimited Roo Code with the Claude APIUnlimited Claude in JetBrains IDEsUnlimited Aider with the Claude APIUnlimited Claude in VS CodeOpenCode config for Claude: opencode.json setup, models and common errorsRunning an Unlimited Hermes Agent on ClaudeRunning Unlimited OpenClaw on ClaudeWhat an Anthropic-Compatible API MeansClaude Code API key — ANTHROPIC_API_KEY: how to get it, set it up, and when to use claude setup-token insteadClaude Code API Costs — Pricing and How to Cut ThemClaude API errors explained: server error 500, overloaded 529, rate limit 429, bad request 400 — and the fix for eachHow to buy Claude API credits — and why a Claude credit beats any "buy Claude account" offerClaude free API: is there a free Claude API key, and what is actually free in 2026?Claude API vs Claude Pro/Max Subscription — Which to ChooseRunning the Codex CLI on ClaudeConnecting Claude to n8nRouting LiteLLM to ClaudeUsing Claude in Kilo CodeBring Claude into GitHub CopilotWire Zed to Claude Through the GatewayConnect Cherry Studio to ClaudeChatAnthropic in LangChain: import, base_url, streaming, tools and errorsClaude with the Vercel AI SDK: createAnthropic baseURL, streaming and fixesRun Factory Droid on ClaudeThe Claude Code API, ExplainedIntegrating the Claude API Into Your ApplicationHow to get unlimited Claude usage — and why there is no "Claude Code cracked"