Claude Opus 5.5 Is Out: Lower Prices, Higher Limits and What Changes for Claude Code
Claude Opus 5.5 launched on September 22, 2026, with lower token prices, higher subscription limits, and changes to Claude Code defaults and API behavior. This guide covers what developers need to know as of September 23: how pricing compares with Opus 5, how to select the model, and which request settings need migration. It also explains why long sessions still need cost monitoring and why progress messages may disappear between tool calls.
Claude Opus 5.5: availability, specifications and pricing
The API model ID is claude-opus-5-5. It has a default 1,000,000-token context window, a maximum output of 128,000 tokens, and always-on adaptive thinking.
Availability includes the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, and Microsoft Foundry. Subscription access covers Pro, Max, Team, and Enterprise. Anthropic describes the model as designed for long-running coding, agents, and professional work; our Claude Opus 5.5 model reference provides a related reference point.
Opus 5.5 vs Opus 5 prices
The launch announcement lists these standard prices, in US dollars per million tokens:
| Token category | Opus 5.5 | Opus 5 | Reduction |
|---|---|---|---|
| Input | $4 | $5 | 20% |
| Output | $20 | $25 | 20% |
| Cache read | $0.20 | $0.50 | 60% |
| Five-minute cache write | $5 | $6.25 | 20% |
For Opus 5.5, one-hour cache writes cost $8 per million tokens. Batch API input and output cost $2 and $10 per million tokens respectively, a 50% discount from standard input and output pricing.
These categories matter separately. A workload with substantial cached input gets a larger reduction on cache reads than on fresh input or output. The total change in its bill therefore depends on the mix of tokens it uses.
What “40% cheaper” means
Anthropic reports that typical workloads at default settings cost 40% less than Opus 5, combining lower prices with fewer tokens per task. That is a workload estimate, not a guaranteed reduction for every request or coding session.
The launch also reports output generation more than 30% faster. Fast mode is a separate option available in Claude Code and the Claude Platform, advertised as up to 2.5× faster. It costs $8 per million input tokens and $40 per million output tokens—twice the standard input and output rates—and the API release notes label it a research preview.
What changes in Claude Code
Claude Code v2.1.280, released September 22, adds Opus 5.5 as the default Opus model. It also changes the default model on Pro and Team Standard from Sonnet to Opus, matching Max, Team Premium, and Enterprise.
That makes the release relevant even if you usually accept the default model. For Pro and Team Standard users, the change includes both the new Opus version and a switch from the Sonnet default.
Update and select the model
Upgrade Claude Code with:
claude update
Select Opus 5.5 inside a session:
/model claude-opus-5-5
Or select it when starting Claude Code:
claude --model claude-opus-5-5
Configuration options also include ANTHROPIC_MODEL, the "model" settings key, and ANTHROPIC_DEFAULT_OPUS_MODEL for the Opus alias. The documented alias change occurs at v2.1.280, so check your installed version when diagnosing an unexpected model selection. The Claude Code model guide is a useful companion for model configuration.
Check effort after upgrading
Opus 5.5 defaults to medium effort, while Opus 5 defaults to high. This matters when comparing cost and behavior: the model’s default effort setting has changed alongside its prices.
Version 2.1.280 also stops effort levels saved before /effort became per-model from automatically carrying over to newly released models. Use /effort to choose a level explicitly, and check any project or managed effortLevel setting that remains relevant.
For a meaningful comparison, record the selected effort as well as the model. Comparing an Opus 5 session at high with an Opus 5.5 session at medium includes an effort-setting difference.
Higher limits: separate five-hour and weekly changes
The September 22 announcement increases five-hour limits for Pro, Max, Team, and seat-based Enterprise. It also announces a subscription rate-limit reset that users can save and use whenever they choose.
The announcement does not quantify the five-hour increase or provide exact allowances for each plan. There is therefore no supported percentage or message count to attach to this launch change.
The weekly increase happened earlier
The Claude Code weekly-limit notice describes a separate adjustment:
| Period | Weekly-limit allowance |
|---|---|
| May 13–September 13, 2026 | Temporary 50% increase over pre-promotion limits |
| From September 14, 2026 | Permanent 25% increase over pre-promotion limits for eligible plans |
The temporary promotion excluded Free and consumption-based Enterprise seats. It did not affect five-hour limits.
The comparison baseline is important. The permanent weekly allowance is higher than the pre-promotion allowance, but lower than the temporary promotional allowance. It is not a new 25% increase on top of the temporary 50% increase.
Use /usage to inspect your limits. When investigating an interruption, distinguish the five-hour window from the weekly allowance; the Claude Code rate-limit guide provides related context.
API migration: settings and response handling to change
Switching the model ID alone may leave incompatible settings in an existing integration. The API release notes document request changes that can return HTTP 400.
Thinking and tool choice
Opus 5.5 uses always-on adaptive thinking. These thinking configurations fail:
thinking: {"type":"disabled"}- Manual-budget thinking using
{"type":"enabled", ...}
Omit thinking or use {"type":"adaptive"}, and control depth through output_config.effort. Disabling thinking is no longer a supported way to reduce spending.
The tool_choice types any and tool also return 400. Use auto, with strict tool use where supported. Review these settings before moving an existing tool workflow to the new model.
Computer use depends on the provider
On the Claude API and Google Cloud, replace computer_20251124 with computer_toolset_20260801; the old tool returns 400.
Amazon Bedrock continues accepting computer_20251124. Apply the migration for the provider your integration actually uses rather than assuming every endpoint has the same requirement.
Preserve thinking through tool loops
The Opus 5.5 migration guide requires thinking blocks to be returned complete and unmodified. Editing, reordering, or partially dropping them can produce 400 errors.
Keep conversations append-only. For accounts created on or after August 31, 2026, at 00:00 UTC, replaying thinking after edits to earlier messages, system, or tools produces 400 by default. Conversation-editing logic therefore needs attention alongside the request settings.
Why progress text can disappear
Narration between tool calls now arrives in thinking blocks rather than text blocks. The default thinking.display: "omitted" leaves the text of those thinking blocks empty.
For visible updates, use "summarized", or use "updates" with the beta header thinking-display-updates-2026-08-18. Render non-empty thinking blocks in the interface. A client that only displays text blocks can otherwise lose the visible progress narration.
Long sessions: caching, thinking and actual consumption
A 1-million-token context window is a capacity specification, not a flat session price. Cached context still incurs charges each time it is read.
At launch rates, reading 1 million cached tokens costs $0.20 on Opus 5.5, compared with $0.50 on Opus 5. Repeating that read 100 times costs $20 versus $50.
Those figures exclude cache writes, fresh input, output, and other charges. They illustrate repeated-read arithmetic; they are not estimates for a complete coding session.
Hidden thinking still costs money
Thinking tokens are charged as output, and max_tokens covers thinking plus response text. Changing thinking visibility does not remove that cost.
Higher effort can increase token consumption. Re-baseline costs at the effort level you intend to use, especially when moving from Opus 5’s high default to Opus 5.5’s medium default. Evaluate fast mode separately because its input and output prices are higher.
Use usage statistics and cache diagnostics
Claude Code’s /usage displays session token statistics and an estimated dollar figure. For Pro and Max, included usage is covered by the subscription, so the displayed estimate is not itself a subscription invoice. The Claude Code pricing guide is a related reference for interpreting costs.
Prompt-cache statistics predate this launch: they require v2.1.251, while likely cache-miss causes were added in v2.1.260.
On September 23, API cache diagnostics became generally available. The cache-diagnosis-2026-04-07 beta header is no longer required; opt in with a diagnostics object on Messages requests. Responses from POST /v1/messages now always include diagnostics, set to null when diagnostics were not requested.
Key takeaways
- Claude Opus 5.5 launched September 22 with model ID
claude-opus-5-5, a default 1-million-token context window, and always-on adaptive thinking. - Standard input and output prices fell 20%; cache-read prices fell 60%. Anthropic’s 40% typical-workload saving is an estimate at default settings.
- Claude Code v2.1.280 adds support, makes Opus 5.5 the default Opus model, and switches Pro and Team Standard defaults from Sonnet to Opus.
- Five-hour limits increased without published per-plan amounts. The permanent 25% weekly increase began September 14 and uses pre-promotion limits as its baseline.
- API migrations need checks for thinking settings, forced tool choice, computer-use tools, preserved thinking blocks, and progress rendering.
- Long sessions still accumulate cache-read and output charges, including hidden thinking. Measure consumption at your chosen effort setting.
FAQ
How much does Claude Opus 5.5 cost compared with Opus 5?
Standard Opus 5.5 pricing is $4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Opus 5. Cache reads cost $0.20 per million tokens instead of $0.50; Anthropic’s separate 40% typical-workload saving combines price reductions with fewer tokens per task at default settings.
Which Claude Code version supports Opus 5.5, and how do I select it?
Claude Code v2.1.280 adds Opus 5.5 support and makes it the default Opus model. Update with claude update, then select it with /model claude-opus-5-5 or start Claude Code with claude --model claude-opus-5-5.
Does the 1-million-token context window make long sessions expensive?
Costs depend on token consumption, including repeated cache reads, cache writes, fresh input, and output. Reading 1 million cached tokens 100 times costs $20 on Opus 5.5, excluding other charges; hidden thinking also counts as billable output.
Can I disable thinking, and why did progress messages disappear?
Opus 5.5 does not accept disabled or manually budgeted thinking; omit the thinking setting or use adaptive thinking, with depth controlled through output_config.effort. Progress narration now arrives in thinking blocks, whose text is empty under the default omitted display setting. Use summarized display, or updates with the required beta header, and render non-empty thinking blocks.
One API key for Claude Opus 5.5, Sonnet 5, Haiku 4.5 and Fable 5.1, plus GPT-6 models. Pay as you go, no subscription.
Get Your API Key →