AI Newsway

Anthropic's Sonnet 5.5 Beats Opus 5.5 on Terminal-Bench 4.0

The mid-tier model holds Sonnet 5's $2 per million input token price while Anthropic claims 30% more speed and up to 30% less cost per task

|5 min read0
AI Summary
Anthropic released Claude Sonnet 5.5 on September 28, 2026, and it scores 70.6% on the Terminal-Bench 4.0 agentic coding evaluation against 66.4% for the pricier Claude Opus 5.5. Pricing is unchanged at $2 per million input tokens and $10 per million output; the claimed 30% cost reduction per task comes from fewer tokens and tool calls. Higher-risk cybersecurity requests now fall back to Sonnet 5, a first for the mid-tier line.
A user working with a conversational AI assistant, the everyday task profile Anthropic says Sonnet 5.5 is tuned for.
A user working with a conversational AI assistant, the everyday task profile Anthropic says Sonnet 5.5 is tuned for.

Anthropic released Claude Sonnet 5.5 on Monday, and the number worth noting is not the price β€” it is a benchmark where the cheap model beats the expensive one. On Anthropic's own scorecard, Sonnet 5.5 posts 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, against 66.4% for Claude Opus 5.5. Sonnet 5, three months old, managed 10.3% on the same test.

Key takeaways

  • Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 versus 66.4% for the more expensive Opus 5.5.
  • Pricing is unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output tokens; the cost saving comes from using fewer tokens, not a lower rate.
  • Balyasny Asset Management reported answers costing 121,000 tokens where Sonnet 5 spent 497,000.

Where the 30% cost cut actually comes from

Anthropic's headline claims are 30%-plus faster generation and up to 30% lower cost per task. The second figure is easy to misread. The sticker price did not move: input stays at $2 per million tokens, output at $10, with cache reads at $0.20 and cache writes at $2.50. Savings come from the model finishing jobs with fewer tokens and fewer tool calls, which is a different kind of win β€” it only shows up on an invoice if your workload is agentic.

The customer figures Anthropic published make that concrete. Slack reported matching or beating Sonnet 5 across nearly all its evaluations in fewer steps and with roughly 14% fewer output tokens. Box measured results 2.4 times faster with 12% fewer total tokens. Lovable said tasks finished with a third fewer tool calls and about half the shell runs. The sharpest number came from Balyasny Asset Management: 121,000 tokens per answer where Sonnet 5 had burned 497,000.

Reading the benchmark table

Sonnet 5.5 lands close to Opus 5.5 nearly everywhere and ahead of it in one place. On OSWorld 2.1 it scores 80.1% to Opus 5.5's 81.8%, up from 57.0% for Sonnet 5. Humanity's Last Exam gives 64.5% against 67.7%. On GDPval-AA v2.1 the two are effectively tied at 1844 and 1846, and on AA-Briefcase v1.1 the gap is 1811 to 1822. Terminal-Bench is also the only one of the three tests Anthropic files under agentic coding that Sonnet 5.5 wins: Opus 5.5 leads on CursorBench 4.0, 57.8% to 55.5%, and on FrontierCode 1.1, 54.4% to 46.2%. Anthropic notes Sonnet 5.5 scores lower at Max effort than at Xhigh on FrontierCode, because at Max it more often ran a code-review skill that produced out-of-scope edits and timeouts, which the benchmark penalises.

The Sonnet 5 column is where caution is warranted. A jump from 10.3% to 70.6% on Terminal-Bench 4.0, or 15.6% to 61.6% on Chartography, says more about how poorly the previous generation handled those specific harnesses than about a sevenfold gain in capability. The like-for-like comparison to read is Sonnet 5.5 against Opus 5.5, and there the story is that a mid-tier model now sits within a couple of points of the flagship on most work β€” while spawning more parallel agents inside the same budget.

What Anthropic put behind a fallback

Sonnet 5.5 arrives with a constraint its predecessors did not have. Anthropic says higher-risk cybersecurity tasks will visibly fall back to Sonnet 5, an explicit downgrade rather than a refusal. TechCrunch reported that this makes Sonnet the first model in its tier to sit under the cyber safeguards previously reserved for Opus and Fable, on the basis that its cyber capability is now comparable to Opus 5. It is also the first Sonnet shipped with classifiers meant to stop rivals extracting its reasoning traces for distillation. Biology safeguards carry over from Sonnet 5 unchanged.

Context: the mid-tier is where the fight is

Nobody is competing on flagship pricing anymore. OpenAI halved its mid-tier token prices with Sol and Luna last week, and Meta announced a model to drive a smart-glasses feature. Anthropic's answer is not a cheaper rate card but a model that consumes less of it, backed by up to 90% savings with prompt caching and 50% with batch processing.

Outlook

Sonnet 5.5 is available on Claude.ai and, for developers, on the Claude Platform as well as Amazon Web Services, Google Cloud and Microsoft Azure, under the identifier claude-sonnet-5-5. A refreshed Haiku is expected within weeks, with no date committed. The open question is what Opus 5.5 is still for: if the cheap model takes the terminal-coding benchmark and ties on knowledge work, the flagship's remaining claim is judgment on problems that are hard to benchmark.

FAQ

Is Claude Sonnet 5.5 more expensive than Sonnet 5?

No. It costs the same: $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 and cache writes at $2.50. Anthropic's claim of up to 30% lower cost per task comes from the model using fewer tokens and tool calls, not from a reduced rate.

Does Sonnet 5.5 really beat Opus 5.5?

On Terminal-Bench 4.0, an agentic coding evaluation, yes β€” 70.6% to 66.4%, and Anthropic reports that Opus 5.5 figure at its best effort setting. It is the only one of the three agentic coding benchmarks Anthropic published where Sonnet 5.5 comes out ahead; Opus 5.5 leads on CursorBench 4.0 and FrontierCode 1.1. Opus 5.5 also stays slightly ahead on OSWorld 2.1 and Humanity's Last Exam, by one to three points.

Where can developers use Sonnet 5.5?

It is available natively on the Claude Platform and through Amazon Web Services, Google Cloud and Microsoft Azure, using the model identifier claude-sonnet-5-5. Consumer access is via Claude.ai.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Opus 5.5 at Max Effort Waited 682 Seconds Before Its First Token
LLM & Chatbots

Opus 5.5 at Max Effort Waited 682 Seconds Before Its First Token

Artificial Analysis measured Claude Opus 5.5's max-effort configuration at 682.71 seconds to first token, against a 3.79-second median for comparable models.

Seung Jung22 hours ago
Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing
LLM & Chatbots

Ox Alpha Unmasked: Z.ai Ships GLM-5.3-Flash With MIT Weights and 15-Cent Pricing

Z.ai confirmed the stealth Ox Alpha model is GLM-5.3-Flash: 320B parameters, 18B active, MIT-licensed weights and sub-dollar output pricing.

Seung Jung32 days ago
Anthropic Gives Claude and Cowork a Single Shared Memory
LLM & Chatbots

Anthropic Gives Claude and Cowork a Single Shared Memory

Anthropic merged Claude and Cowork memory into one shared store, on by default for Free, Pro and Max users, with no option to keep the two products apart.

Seung Jung34 days ago
Anthropic Warns Infostealer Malware Is Draining Paid Claude Accounts
LLM & Chatbots

Anthropic Warns Infostealer Malware Is Draining Paid Claude Accounts

Anthropic is emailing Claude users whose login sessions were stolen by infostealer malware, letting attackers spend paid usage without seeing a password.

Seung Jung19 days ago
Mercury 2.5 Ships at 1,107 Tokens a Second and Four Cents a Million
LLM & Chatbots

Mercury 2.5 Ships at 1,107 Tokens a Second and Four Cents a Million

A phone agent startup says Mercury cut its worst-case response from minutes to one second. Inception Labs' new diffusion model is built around that kind of number.

Seung Jung15 days ago
Claude Is Adults-Only, and Its Age Detector Keeps Locking Out Adults
LLM & Chatbots

Claude Is Adults-Only, and Its Age Detector Keeps Locking Out Adults

Anthropic published the mechanics of Claude's 18+ enforcement, including classifier-driven suspensions and a Yoti appeal route for wrongly flagged adults.

Seung Jung17 days ago