Claude Sonnet 5.5 Launches: Pricing, What's New and Haiku 5.5 on the Way
Claude Sonnet 5.5 launched on September 28, 2026, with unchanged token pricing and improvements that Anthropic reports in speed, task cost and coding benchmarks. This guide covers pricing, the differences from Sonnet 5, API settings to check before migrating, and what is known about Haiku 5.5 as of September 29.
Claude Sonnet 5.5: launch, availability and limits
Sonnet 5.5 is the second model in the Claude 5.5 family, following Opus 5.5. Anthropic’s launch announcement positions it for everyday work, including bug fixing, document creation, slides, spreadsheets, design, clearer writing and collaboration.
The model is available through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. Its Claude API identifier is claude-sonnet-5-5.
For developers, the main reference points are:
| Specification | Claude Sonnet 5.5 |
|---|---|
| Launch date | September 28, 2026 |
| Claude API model ID | claude-sonnet-5-5 |
| Context window | 1 million tokens |
| Maximum output | 128,000 tokens |
| Supported input | Text and images |
| Supported output | Text |
The context window and maximum output are separate limits. When planning a request, check both the material you want the model to process and the response length your application needs.
The Sonnet 5.5 model reference is a useful companion for model selection. For Claude Code workflows, the Claude Code models guide provides a separate place to review model choices.
Sonnet 5.5 pricing: unchanged rates, lower reported task costs
Standard Sonnet 5.5 pricing remains $2 per million input tokens and $10 per million output tokens, matching Sonnet 5. Cache reads cost $0.20 per million tokens.
The official model overview also lists cache-write rates and the Batch API discount:
| Token category | Price per million tokens |
|---|---|
| Standard input | $2.00 |
| Standard output | $10.00 |
| Cache read | $0.20 |
| Five-minute cache write | $2.50 |
| One-hour cache write | $4.00 |
The Batch API provides a 50% discount on input and output. Keep that discount distinct from the standard rates and the separate prompt-caching categories when estimating a workload’s cost.
What “up to 30% less” means
Anthropic reports task costs of “up to 30% less” in its testing. That claim concerns fewer tokens used to complete a task; it does not describe a reduction in the price of each token.
This distinction matters when comparing sonnet 5.5 pricing with Sonnet 5. If a task consumes the same quantities of standard input and output tokens on both models, the listed rates provide no per-token saving. The reported saving depends on how much work Sonnet 5.5 completes with fewer tokens.
Treat the figure as a workload-dependent result, rather than a forecast for every application. A practical comparison should track completed tasks, token consumption and resulting cost together. Looking only at the rate card would miss the efficiency claim; assuming a universal 30% saving would overstate it.
Prompt caching in a cost estimate
Cache reads, five-minute writes and one-hour writes have different prices. An estimate that includes caching should therefore distinguish those categories from standard input and output.
For a broader discussion of token budgets, see the guide to reducing Claude token usage. For this launch, the central pricing point is straightforward: token rates are unchanged, while Anthropic reports lower token consumption per task in its testing.
Sonnet 5.5 vs Sonnet 5: what improved?
Anthropic reports output generation that is “30%+ faster” than Sonnet 5. It also publishes gains across terminal work, coding and professional-task evaluations.
| Anthropic-published evaluation | Sonnet 5 | Sonnet 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 10.3% | 70.6% |
| CursorBench 4.0 | 34.1% | 55.5% |
| GDPval-AA v2.1 | 1449 | 1844 |
These are Anthropic’s evaluation results. They provide evidence for the company’s launch claims, but they should be read alongside the task types your own application handles.
The speed figure specifically concerns output generation. Avoid treating it as a promise that every complete workflow will finish 30% faster. Likewise, the task-cost figure describes results from Anthropic’s testing and depends on the workload.
Where Opus 5.5 still fits
Sonnet 5.5’s GDPval-AA v2.1 score of 1844 is close to Opus 5.5’s 1846 on that evaluation. That proximity is useful context, but it does not establish equivalent performance across all work.
Anthropic still describes Opus 5.5 as stronger for complex, open-ended tasks that require sustained judgment. Sonnet 5.5’s positioning emphasizes everyday tasks and collaboration, including bug fixing and creating documents, slides and spreadsheets.
For a migration evaluation, choose examples that reflect those distinctions. Test routine bug fixes separately from work that requires extended judgment, and assess whether the resulting output meets your requirements. A single aggregate benchmark cannot answer every model-selection question.
The tokenizer has not changed
Sonnet 5.5 uses the same tokenizer as Sonnet 5: identical text produces identical token counts.
That makes the reported efficiency improvement easier to interpret. A fixed piece of text does not become cheaper because the new model counts it differently. The launch’s lower-task-cost claim concerns fewer tokens per completed task.
Effort and thinking defaults to check
Sonnet 5.5 defaults to Medium effort in Claude Code and Claude apps, and High effort on the Claude Platform. Adaptive thinking is enabled by default.
Those defaults matter when comparing results across environments. If you evaluate the model in Claude Code and then through the API, record the effort setting alongside the model identifier so that the comparison is easier to interpret.
The Sonnet 5.5 changes documentation identifies output_config.effort as the API control for effort.
Disabling up-front thinking
The migration instructions replace thinking: {"type": "disabled"} with the following setting to disable up-front thinking at high effort or below:
{
"thinking": {
"type": "between_tools"
}
}
The scope of that setting is important: it disables up-front thinking under the stated effort condition. Preserve that distinction when updating configuration or explaining the change to other developers.
If your application currently sends the older disabled setting, include it in the migration review. Changing only the model ID would leave that configuration unchecked.
API migration: tools, advisors and preserved thinking
The September 28 release notes identify several compatibility changes relevant to existing integrations. Review these before moving a Sonnet 5 workload to claude-sonnet-5-5.
Forced tool choices return HTTP 400
Sonnet 5.5 rejects tool_choice types any and tool with HTTP 400. The supported alternatives are auto and none.
If a workflow currently uses either rejected type, replacing the model identifier alone is insufficient. Review the request configuration and the application behavior that depends on it.
The supported alternatives should not be treated as direct equivalents to a forced-tool request. Check whether the revised configuration still meets the workflow’s requirements before completing the migration.
Computer-use tooling changes
On the Claude API and Google Cloud, Sonnet 5.5 rejects computer_20251124. The documented replacement is computer_toolset_20260801.
Apply that change within the stated platform scope. The release information names the Claude API and Google Cloud; it does not provide a basis for assuming identical migration details on every listed provider.
Advisor model restrictions
The advisor tool rejects Claude Opus 4.8, Claude Opus 4.7 and Claude Sonnet 5 as advisors.
Check advisor configuration separately from the primary model selection. If one of those models appears in an advisor role, address that dependency as part of the migration. The launch brief does not supply a replacement advisor list, so no substitute should be inferred from these restrictions alone.
Preserved thinking has account boundaries
Thinking blocks are tied to their model and conversation. Sonnet 5.5 adds an account constraint: its blocks work only in the account that produced them or a linked account.
Blocks submitted from another account are dropped while the request succeeds. This means a successful request alone does not confirm that all submitted thinking blocks were preserved.
If your application reuses conversation material across accounts, include this behavior in your review. Check model, conversation and account compatibility together when deciding which thinking blocks to carry forward.
Haiku 5.5 is announced, with release details still pending
Anthropic announced Claude Haiku 5.5 alongside the Sonnet launch and said it would join the family “in the coming weeks.” Its stated focus is high-volume, cost-sensitive applications.
As of September 29, 2026, Haiku 5.5 has been announced but has not been released. The September 28 announcement supplied no exact launch date and no pricing.
That leaves two practical limits for planning. There is no announced date to use as a firm migration deadline, and there is no published price to use in a Sonnet-versus-Haiku cost comparison.
Developers can identify workloads they may want to evaluate once Haiku 5.5 arrives, particularly high-volume applications. Any decision that depends on its price or release timing must wait for those details.
Key takeaways
- Claude Sonnet 5.5 launched September 28, 2026, with the Claude API identifier
claude-sonnet-5-5. - Standard pricing matches Sonnet 5: $2 per million input tokens, $10 per million output tokens and $0.20 per million cache-read tokens.
- Anthropic reports faster generation and lower task costs, with “30%+ faster” output generation and tasks costing “up to 30% less” in its testing.
- Migration requires configuration checks, including thinking settings, forced tool choices, computer-use tooling, advisor models and preserved thinking.
- Haiku 5.5 is forthcoming, but its exact launch date and pricing were not announced by September 29.
FAQ
When did Claude Sonnet 5.5 launch, and what are its limits?
Claude Sonnet 5.5 launched on September 28, 2026, as the second model in the Claude 5.5 family. It supports a 1-million-token context window and a maximum output of 128,000 tokens, with text and image input and text output.
How much does Sonnet 5.5 cost, including prompt caching?
Standard pricing is $2 per million input tokens and $10 per million output tokens; cache reads cost $0.20 per million tokens. Five-minute cache writes cost $2.50 per million tokens, one-hour cache writes cost $4, and the Batch API provides a 50% discount on input and output.
Which API settings need attention when migrating to Sonnet 5.5?
To disable up-front thinking at high effort or below, replace thinking: {"type": "disabled"} with thinking: {"type": "between_tools"}. Forced tool_choice types any and tool return HTTP 400, while auto and none are supported. Also review computer-use tooling, advisor-model restrictions and the account requirements for preserved thinking blocks.
When will Haiku 5.5 launch, and has pricing been announced?
Anthropic said Haiku 5.5 would arrive “in the coming weeks” for high-volume, cost-sensitive applications. As of September 29, 2026, it had not been released, and the announcement supplied neither an exact launch date nor pricing.
One API key for Claude Opus 5.5, Sonnet 5, Haiku 4.5 and Fable 5.1, plus GPT-6 models. Pay as you go, no subscription.
Get Your API Key →