GPT-6 Sol & Claude Opus 5.5 are liveGPT-6 Sol at half the price of 5.6 Sol
Blog
Guide

Claude Sonnet 5.5 review

A faster daily model with a more important API migration than the version number suggests.

12 min readOmniaKey
Claude Sonnet 5.5Model reviewAPI pricingClaude CodeMigration

Claude Sonnet 5.5 is the practical upgrade in Anthropic’s Claude 5.5 family. It launched on September 28, 2026 with the same standard API price as Sonnet 5, but Anthropic says it generates output 30% or more quickly and can cost up to 30% less per task because it uses fewer tokens. Its published coding and knowledge-work scores also move much closer to Opus 5.5.

That does not make Sonnet 5.5 a universal replacement for Opus. The model is strongest when the work is well scoped: daily coding, bug fixes, document production, spreadsheet analysis and fast agent iteration. The bigger surprise for developers is the API contract. A Sonnet 5 integration that relies on disabled thinking, forced tool calls, old computer-use tools or untyped response handling needs a migration pass before the model name is changed.

Checked September 29, 2026. This review uses Anthropic’s launch announcement and current Claude Platform documentation. The benchmark scores and speed/cost claims below are Anthropic-reported; we did not run an independent paid benchmark or latency test. Cost examples use the direct global API price and hold token counts constant.

Claude Sonnet 5.5 at a glance

SpecificationClaude Sonnet 5.5
Release dateSeptember 28, 2026
Claude API model IDclaude-sonnet-5-5
Context window1M tokens
Standard maximum output128K tokens
Input / output$2 / $10 per million tokens
Cache write / read$2.50 / $0.20 per million tokens
ThinkingAdaptive by default
Default efforthigh
Comparative latencyFast
Knowledge cutoffJune 2026
Input and outputText and images → text

The model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The platform-specific model ID can differ; check the provider’s availability table before changing a production client. The standard API has a 1M-token context window and 128K maximum output. Batch API supports up to 300K output tokens with its documented beta header, which is a separate limit from synchronous Messages requests.

What improved over Sonnet 5?

Anthropic’s headline claim is 30%+ faster generation. The company also says Sonnet 5.5 typically costs up to 30% less per task, even though its published token rates are unchanged. The explanation is lower token consumption and fewer tool steps, rather than a price cut on each token.

The launch material reports these results against Sonnet 5:

EvaluationSonnet 5.5Sonnet 5What it measures
Terminal-Bench 4.070.6%10.3%Multi-step terminal coding
CursorBench 4.055.5%34.1%Ambiguous coding-agent tasks
GDPval-AA v2.118441449Knowledge work across occupations
AA-Briefcase v1.118111359Long-horizon knowledge work
Humanity’s Last Exam64.5%54.9%Multidisciplinary reasoning with tools
OSWorld 2.180.1% partial57.0% partialComputer use
Chartography61.6%15.6%Visual chart recognition

These are useful signals, not a promise about your repository. Sonnet 5.5’s CursorBench score is close to Opus 5.5’s 57.8%, and its GDPval score is close to Opus 5.5’s 1846. Anthropic still describes Opus as clearly stronger for complex, open-ended work that requires sustained judgment. The sensible reading is that Sonnet 5.5 closes the gap on many repeatable tasks; it does not erase the reason to keep an escalation model.

Anthropic’s cost charts are also easy to misread. They compare each model across effort levels and count the cost of completing a task, not just a single response. Sonnet 5.5 can reach a higher score with fewer tokens on some workloads, but the result depends on effort, tools, prompt length, retries and the acceptance test. Keep those variables fixed when you compare it with the Opus 5.5 model page or your current Sonnet route.

API pricing: the rate card did not change

Sonnet 5.5 keeps Sonnet 5’s direct Anthropic rates:

Billing itemPrice
Uncached input$2 / MTok
Output, including billed thinking tokens$10 / MTok
Five-minute cache write$2.50 / MTok
One-hour cache write$4 / MTok
Cache read$0.20 / MTok
Batch API input and output50% discount

For a simple uncached task with 100K input tokens and 20K output tokens, the direct API cost is 0.1 × $2 + 0.02 × $10 = $0.40. If 80K of the input is a cache hit, the same token counts cost 0.02 × $2 + 0.08 × $0.20 + 0.02 × $10 = $0.376, before the cache-write charge. Real agent bills vary because thinking tokens, tool calls, retries and the number of turns change with the task.

This is why “up to 30% less per task” should not be read as “30% off every invoice.” Measure cost per accepted change: total input, cache, output and tool-related charges divided by work that passed your tests or review. OmniaKey’s gateway quote is a separate price surface; check the current model catalog before budgeting a route.

The migration details that can break a working client

The model ID is simple: replace the old ID with claude-sonnet-5-5. The request settings are not always interchangeable.

  1. Thinking is adaptive by default. Omit the thinking field or send {"type":"adaptive"}. To turn off up-front thinking, use thinking: {"type":"between_tools"}. The old {"type":"disabled"} value returns a 400 error, and manual budget_tokens are not accepted.
  2. Forced tool use is unsupported. tool_choice: {"type":"any"} and {"type":"tool"} return 400. Use auto or none; if the arguments must match a schema, keep auto and use strict tool definitions or structured outputs.
  3. Do not assume thinking blocks are portable. They are bound to the model and conversation. Sonnet 5.5 can read Sonnet 5 thinking blocks, but Opus 5 and Opus 5.5 cannot read blocks produced by Sonnet 5.5. A mid-conversation model switch can therefore change what the next request sees.
  4. Preserve the conversation before editing it. For accounts subject to the newer binding checks, changing an earlier system message, tool definition or message before a thinking block and then replaying that block can return 400. Append new instructions or follow the documented block-binding controls.
  5. Update computer use on the right platforms. Claude API and Google Cloud require computer_toolset_20260801; the older computer_20251124 tool remains accepted on Amazon Bedrock. This is a tool-loop change, not just a renamed model.
  6. Read streamed responses by block type. Text between tool calls may arrive as thinking progress-update blocks. A client that only renders text blocks can appear to go silent even though the request succeeded.

A minimal adaptive request looks like this:

json
{
  "model": "claude-sonnet-5-5",
  "max_tokens": 4096,
  "output_config": {"effort": "high"},
  "messages": [
    {"role": "user", "content": "Review this migration and list the three largest risks."}
  ]
}

Do not carry over non-default temperature, top_p or top_k values without checking the new model docs; the current Sonnet 5.5 endpoint rejects them. max_tokens must leave room for both thinking and visible output.

Where Sonnet 5.5 fits in a real workflow

Start with Sonnet 5.5 for a scoped feature, a reproduced bug, tests around a known interface, a focused code review, a document or slide draft, and high-volume agent steps. It is especially attractive when the repository has deterministic checks: tests, type checks, builds, screenshots or a reviewer can quickly reject a weak patch.

Use a higher-effort setting or a stronger model when the task is still underspecified, crosses several services, or needs a difficult architectural decision. If Sonnet had the right files and tools, tried seriously and still misunderstood the root cause, that is a capability signal. If it skipped verification because the request was vague or the environment was incomplete, improve the context before paying for a more expensive model.

For Claude Code, keep the model and acceptance test explicit. Compare a fresh Sonnet 5.5 session with the same repository state, tool permissions and prompt used for Sonnet 5. Record first useful output, total elapsed time, billed tokens, retries, failed tool calls and human repair minutes. A model that is 30% faster per token can still lose if it creates one extra review cycle.

Verdict

Claude Sonnet 5.5 is the new default worth testing for everyday coding and knowledge work. The combination of unchanged token rates, lower reported task cost, faster generation and a much higher Sonnet 5 baseline makes the upgrade easy to justify for bounded work.

The upgrade is not a blind model-name swap. Test the thinking mode, tool choice, response-block parser and computer-use path before production. Keep Opus 5.5 for ambiguous architecture, difficult debugging and high-consequence judgment, then let accepted-task data decide where Sonnet 5.5 is sufficient. Read the Opus 5.5 vs Sonnet 5 comparison when the choice is between the two rather than between Sonnet 5 and 5.5.

Frequently asked questions

When was Claude Sonnet 5.5 released?

Anthropic released Claude Sonnet 5.5 on September 28, 2026. Its direct Claude API ID is claude-sonnet-5-5.

Is Sonnet 5.5 cheaper than Sonnet 5?

The listed per-token prices are the same: $2 per million input tokens and $10 per million output tokens, with $0.20 per million cache reads. Anthropic reports up to 30% lower cost per task because the new model often uses fewer tokens and tool steps.

Is Claude Sonnet 5.5 better than Opus 5.5?

No general ranking follows from the launch data. Sonnet 5.5 is faster and cheaper for well-scoped work; Opus 5.5 remains the stronger choice for difficult, open-ended judgment. Compare accepted results, total cost and correction time on your tasks.

Can I disable thinking on Sonnet 5.5?

You cannot use the old thinking: {"type":"disabled"} setting. Use adaptive thinking, or use thinking: {"type":"between_tools"} when you want to turn off up-front thinking at a supported effort level.

Does Sonnet 5.5 have a 1M context window?

Yes. It lists a 1M-token context window and 128K standard maximum output. Those are capacity limits, not included usage or unlimited credits.

Sources checked