Anthropic Python SDK: install, set base_url, stream, and handle errors
The anthropic package is Anthropic's official Python client for the Claude API. It handles authentication, retries, streaming and typed errors, and it works unchanged against any Anthropic-compatible endpoint: set base_url (or ANTHROPIC_BASE_URL) and the rest of your code stays the same.
Install
pip install anthropicThe 1.x releases run on httpx2 and need Python 3.10 or newer. If you pass your own HTTP client, use the SDK's DefaultHttpxClient so its default timeouts and connection limits are kept.
Create a client
import anthropic
client = anthropic.Anthropic(
api_key="your_key", # default: ANTHROPIC_API_KEY
base_url="https://aiprimetech.io", # default: ANTHROPIC_BASE_URL, else https://api.anthropic.com
)base_url is the host root. The SDK appends /v1/messages itself, so https://aiprimetech.io/v1 would produce /v1/v1/messages and a 404 page not found. With both variables set in the environment, anthropic.Anthropic() needs no arguments at all:
export ANTHROPIC_API_KEY=your_key
export ANTHROPIC_BASE_URL=https://aiprimetech.ioSend a message
message = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Summarise TCP slow start in two sentences."}],
)
for block in message.content:
if block.type == "text":
print(block.text)
print(message.usage.input_tokens, message.usage.output_tokens)
print(message._request_id) # log this with any failure reportcontent is a list of blocks (text, thinking, tool use), so check block.type instead of reading content[0].text blindly.
Stream responses
with client.messages.stream(
model="claude-sonnet-5",
max_tokens=16000,
messages=[{"role": "user", "content": "Write a short story about a lighthouse."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()Stream anything with long input, long output or a large max_tokens: a single non-streaming request has to finish inside the timeout, a stream does not. messages.create(..., stream=True) gives you the raw event iterator if you don't need the helper.
Async client
import asyncio, anthropic
client = anthropic.AsyncAnthropic(base_url="https://aiprimetech.io")
limit = asyncio.Semaphore(3) # stay within your key's concurrency
async def ask(question: str) -> str:
async with limit:
msg = await client.messages.create(
model="claude-haiku-4-5",
max_tokens=512,
messages=[{"role": "user", "content": question}],
)
return next(b.text for b in msg.content if b.type == "text")
async def main():
answers = await asyncio.gather(*(ask(q) for q in ["What is DNS?", "What is BGP?"]))
print(answers)
asyncio.run(main())The semaphore matters. Firing hundreds of requests at once with asyncio.gather is the most common cause of Concurrency limit exceeded errors we see from Python clients.
Timeouts and retries
Defaults: a 10-minute timeout per request and 2 automatic retries with exponential backoff. The SDK retries connection errors, 408, 409, 429 and 5xx responses. Because timeouts are retried too, the worst-case wall-clock time is the timeout times the number of attempts.
client = anthropic.Anthropic(base_url="https://aiprimetech.io", timeout=60.0, max_retries=4)
# override for one call only
client.with_options(timeout=300.0).messages.create(
model="claude-opus-5-5",
max_tokens=4000,
messages=[{"role": "user", "content": "Review this design doc..."}],
)Error handling
| Status | Exception | Retry? |
|---|---|---|
| 400 | anthropic.BadRequestError | No: fix the request |
| 401 | anthropic.AuthenticationError | No: wrong or missing key |
| 403 | anthropic.PermissionDeniedError | No |
| 404 | anthropic.NotFoundError | No: model ID or URL |
| 429 | anthropic.RateLimitError | Yes, after a pause |
| 5xx, including 529 overloaded | anthropic.InternalServerError | Yes |
| Network failure | anthropic.APIConnectionError | Yes |
| Timeout | anthropic.APITimeoutError | Yes, or stream |
try:
message = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": "Hi"}])
except anthropic.NotFoundError as e:
print("Bad model ID or URL:", e.message)
except anthropic.RateLimitError as e:
print("Rate limited, retry after", e.response.headers.get("retry-after"))
except anthropic.APIStatusError as e:
print(e.status_code, e.type, e.message)
except anthropic.APIConnectionError:
print("Network problem")Catch the specific classes first and APIStatusError last; one broad except hides the difference between errors worth retrying and errors that need a code change. e.type gives the API's error type string, such as rate_limit_error or overloaded_error.
Model IDs
| Model ID | Use it for |
|---|---|
claude-sonnet-5 | Default for coding and agent work: fast, strong, mid-priced. |
claude-opus-5-5 | Harder reasoning, large refactors, long agent runs. |
claude-fable-5-1 | Most capable and most expensive; save it for the hardest tasks. |
claude-haiku-4-5 | Cheap and quick: titles, summaries, classification, small steps. |
client.models.list() returns the models an endpoint serves, so you can check an ID before you use it. Use the IDs exactly as written: no date suffixes.
Troubleshooting on a gateway
| What you see | Cause | Fix |
|---|---|---|
NotFoundError with 404 page not found | /v1 in base_url | Use the host root |
AuthenticationError | Key not set in the process, or pasted with whitespace | Print len(os.environ["ANTHROPIC_API_KEY"]) |
RateLimitError: Concurrency limit exceeded for user | Too many requests in flight | An asyncio.Semaphore as above |
APITimeoutError on long jobs | Non-streaming request running past the timeout | Stream, or raise timeout |
InternalServerError: No available accounts | Temporary upstream capacity | Let the SDK retry; raise max_retries |
Insufficient account balance | Balance empty | Top up |
More on model-side errors: the Claude API errors guide.
Frequently asked questions
How do I install the Anthropic Python SDK?
How do I set a custom base URL in the Anthropic Python SDK?
What are the default timeout and retries?
How do I stream with the Python SDK?
Get an API key — no Anthropic account or waitlist required.
Get your API keyAI Prime Tech is an independent API gateway. It is not affiliated with, endorsed by, or a reseller of Anthropic. Claude and related model names are trademarks of their respective owners.