Opus 5.5 vs Sonnet 5.5
Choose by results and total cost.
Claude Opus 5.5 vs Sonnet 5.5 is a choice about completed work and its total cost. Start an evaluation with Sonnet for clearly scoped coding, frequent iterations and tasks with dependable acceptance checks. Evaluate Opus when diagnosis remains ambiguous, a change crosses architectural boundaries or a plausible mistake would be expensive to repair. These are workload recommendations, not a universal performance ranking.
Sonnet's standard input and output token rates are half Opus's, but their cache-read price is identical. Meanwhile, Sonnet scores higher on one published coding benchmark and lower on others. Neither “always use the larger model” nor “the cheaper model has replaced it” follows from that evidence.
Sources checked September 29, 2026. This comparison analyzes official documentation and attributed benchmark results. We have not run an independent head-to-head test. The cost examples hold token quantities constant; they are calculations, not measured task bills. OmniaKey provides an API gateway and publishes this guide; its route prices are separate from Anthropic's direct API rates.
Opus 5.5 and Sonnet 5.5 at a glance
Anthropic released Opus 5.5 on September 22 and Sonnet 5.5 on September 28, 2026. This article compares those exact versions. The Opus 5.5 vs Sonnet 5 comparison retains the earlier pairing.
| Specification | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Claude API model ID | claude-sonnet-5-5 | claude-opus-5-5 |
| Context / standard maximum output | 1M / 128K tokens | 1M / 128K tokens |
| Input → output | Text and images → text | Text and images → text |
| API default effort | high | medium |
| Thinking | Adaptive; between_tools available | Adaptive, always on |
| Vendor latency category | Fast | Moderate |
| Standard input / output per 1M tokens | $2 / $10 | $4 / $20 |
| Cache read per 1M tokens | $0.20 | $0.20 |
Source: Anthropic's model overview. Equal context capacity does not establish equal retrieval accuracy or reasoning quality. Input and output must fit the applicable context budget; the 128K output limit is not extra free usage. The separate 300K output option is a Batch API beta, not the normal synchronous limit.
Coding benchmarks: why there is no single winner
The following figures come from the Sonnet 5.5 announcement, which places both models in the same comparison table. They are published results, not an OmniaKey rerun.
| Evaluation | Sonnet 5.5 | Opus 5.5 | What to keep attached to the score |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | Sonnet max; Opus xhigh |
| FrontierCode 1.1 Main | 46.2% at max; 52.1% at xhigh | 54.4% | Scope violations and timeouts can lower scores |
| CursorBench 4.0 | 55.5% | 57.8% | A particular coding evaluation, not all repository work |
| GDPval-AA v2.1 | 1844 | 1846 | Scores, not percentages or measured task success rates |
The Sonnet system card, section 8.5, describes Terminal-Bench as 66 tasks with five trials each, or 330 trials per model. It reports standard errors of ±2.5 points for Sonnet and ±2.6 for Opus, treating trials as independent. Production safeguards were enabled and some requests used fallback models; the affected trial shares differed. The run used Claude Code in --bare mode. A 4.2-point gap therefore should not become “Sonnet is better at every coding task,” or an unsupported claim of statistical significance.
FrontierCode offers a different lesson: more effort does not guarantee a better outcome. Anthropic's footnote says Sonnet at max more often invoked a review workflow that, in examined cases, contributed to a timeout or extra edits outside the requested scope. Report both Sonnet settings rather than selecting whichever supports a preferred winner.
GDPval's two-point difference should not become a claim that the models are interchangeable. The announcement also discloses a pre-release structured-output bug affecting the Sonnet deployment used for GDPval-AA and AA-Briefcase; it was subsequently fixed. Keep the version and conditions beside the figures. For individual release and migration details, see the Sonnet 5.5 review and Opus 5.5 review.
API pricing: is Sonnet really half the cost?
These are Anthropic's standard global direct API rates in USD per million tokens, checked against its pricing documentation. Tools, taxes, regional modifiers and other platforms' charges are separate.
| Token category | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Uncached input | $2 | $4 |
| Billed output | $10 | $20 |
| 5-minute cache write | $2.50 | $5 |
| 1-hour cache write | $4 | $8 |
| Cache read | $0.20 | $0.20 |
Output includes billed thinking, not just the visible answer. Both models have a minimum cacheable prompt length of 512 tokens, but meeting that threshold does not prove a cache hit. Read the provider's usage fields.
Three deliberately fixed workloads show why price per token and cost per task differ. Quantities can be cumulative across requests; each request must still obey context and output limits.
| Same token workload | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| 100K uncached input + 20K billed output | $0.40 | $0.80 |
| The above plus 800K initial 5-minute cache-write tokens | $2.40 | $4.80 |
| 100K uncached + 800K valid cache-read tokens + 20K billed output | $0.56 | $0.96 |
For the last row, Sonnet costs 0.1 × 2 + 0.8 × 0.20 + 0.02 × 10 = $0.56; Opus costs 0.1 × 4 + 0.8 × 0.20 + 0.02 × 20 = $0.96. That is about 1.71 times, not twice, the read-stage bill. The earlier write is additional; the middle row illustrates its cost. Do not count cache-write tokens again as uncached input.
Real tasks can use different token totals, tool calls and retries on each model. The useful measurement is all billed cost, including failed attempts, divided by accepted results. Record human correction time separately before assigning it an explicit hourly value.
Speed and effort: compare the complete task
Anthropic's “30%+ faster output” and “up to 30% lower task cost” claims for Sonnet 5.5 compare it with Sonnet 5, not Opus 5.5. Its API input and output rates did not fall by 30%. Similarly, the roughly 40% typical-task saving announced for Opus 5.5 compares it with Opus 5.
Sonnet defaults to high on the Claude API but to medium in Claude Code and the Claude apps. Opus's API default is medium. Specify the client and effort when comparing results. Even identical effort names do not promise identical compute budgets.
For interactive work, measure time to useful output and total completion time separately. A quick first response can still lead to several repair cycles; a slower successful run may finish the task earlier. The vendor labels “Fast” and “Moderate” are not a measured latency ratio for your region or gateway.
When is Opus 5.5 worth the extra cost?
Start a Sonnet evaluation with work you can check cheaply: a reproduced bug, a small feature with acceptance tests, a focused review or structured document extraction. Its lower token rates matter when many runs meet the same quality bar.
Evaluate Opus when the difficult part is deciding what is wrong or what should change: ambiguous failures across services, migrations with subtle compatibility constraints or long investigations. Keep the same evidence and acceptance checks. Higher price alone does not make a result more correct or safe.
For long documents, test evidence use. Both advertise 1M context, so ask for traceable citations and check omissions instead of choosing by the window size. For visual work, score the actual screenshot or chart task; neither model's image input makes it a native image generator.
This is our cost-aware evaluation strategy. Anthropic's general model guide instead recommends starting with Opus 5.5 for most workloads. A team that prioritizes quality over token spend can reasonably begin there. What matters is disclosing the objective and checking accepted outcomes. The Claude Code model guide covers the broader model-selection workflow.
Switching models requires more than replacing the ID
Read both the Sonnet migration guide and Opus migration guide.
- Sonnet's
between_toolssetting turns off up-front thinking athigheffort or below. It does not mean every form of thinking is disabled. Opus 5.5 keeps adaptive thinking on. - Both reject forced
tool_choicetypesanyandtool. Strict tool input validation does not force a tool call. Check platform support before reusing a request. - Thinking blocks have model, account and conversation-preservation rules. Start a fresh conversation for an isolated comparison; do not edit signed history or transfer it between models by assumption.
- Progress between tool calls can appear in thinking blocks and be omitted from display by default. Read blocks by type and verify streaming behavior before interpreting silence as a stalled agent.
These checks matter even when both models expose the same headline context capacity. A model-selection article cannot establish compatibility for every SDK or gateway.
A reproducible evaluation for your own work
Choose tasks before seeing the results. A practical design is six tasks in each of four groups: bounded coding, difficult debugging/refactoring, tool workflows and document/visual work. With 24 tasks, two models and three independent repetitions, that is 144 task runs, potentially many more API requests. This is a proposed design, not a test we performed.
Use the same repository commit, prompts, tools, region, permissions, step/time limits and acceptance tests. Start each run in a fresh conversation and reset workspace; alternate model order. Compare both at explicit medium as one track, and report any default-setting or higher-effort track separately.
Record accepted results, failures, retries, tool errors, refusals, observed fallbacks, billed tokens, elapsed time and actual human corrections. Include failed attempts in cost. Use hidden tests for code and source checks for documents; do not let a model's own “done” message determine success. Three repetitions per task are insufficient for a precise per-task P95 estimate. If no run succeeds, report no accepted result rather than a zero cost per success.
Frequently asked questions
Is Sonnet 5.5 better than Opus 5.5 for coding?
It scores higher on the published Terminal-Bench 4.0 comparison, while Opus scores higher on the listed FrontierCode and CursorBench results. Settings and uncertainty matter. Choose using representative tasks, not one benchmark headline.
Does Sonnet 5.5 always cost half as much?
Its standard uncached input and output token rates are half Opus's. Cache reads cost the same, and token use, retries, tools and corrections change the full bill.
Which should I use in Claude Code?
Evaluate Sonnet for scoped work and Opus for difficult investigation, or start with Opus if quality is your first constraint. Record effort explicitly. A Claude subscription and metered API billing are separate; the Claude Code pricing guide explains that boundary.
Do I get a larger context window with Opus?
Not in this pair: both list 1M context and 128K standard maximum output. More useful reasoning within the window has to be established by the task.
Current routes and official sources
Check the current Sonnet 5.5 and Opus 5.5 pages, or the model catalog, for OmniaKey's displayed route and price. The direct API calculations above are not a gateway quote or a verified live API call.
The linked Anthropic announcements, specifications, pricing, migration guides and system card are the official English sources used in this comparison. Recheck prices and request behavior before changing a production integration.