Anthropic releases Claude Sonnet 5.5: Terminal-Bench jumps from 10.3% to 70.6% with no price increase

2026-09-30·10 min read

Start with what can be checked. Per Anthropic's official launch page dated September 28, Claude Sonnet 5.5 is the second model in the Claude 5.5 family, a clear upgrade over Claude Sonnet 5, running 30%+ faster and costing up to 30% less for most work. Pricing is stated plainly: the same as Sonnet 5 at 2 dollars per million input tokens, 10 dollars per million output tokens and 0.20 dollars per million tokens for cache reads, but because it typically needs far fewer tokens to do the same work, Anthropic's testing found it costs up to 30% less per task than its predecessor. Reuters added context the same day: this is Anthropic's second model launch as it builds toward a planned IPO, and the company says enterprise customers account for about 80% of its business. This article sticks to what can be checked in the official launch page, the system card and authoritative media reports.

The most striking numbers are in agentic terminal coding. On Terminal-Bench 4.0, Anthropic reports Sonnet 5.5 at 70.6% against Sonnet 5's 10.3%, close to a sevenfold gap; for reference, the same table puts Opus 5.5 at 66.4% at Xhigh effort. On CursorBench 4.0, which evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions, Sonnet 5.5 scores 55.5% versus Sonnet 5's 34.1% and Opus 5.5's 57.8%. On FrontierCode 1.1 main set, Sonnet 5.5 scores 46.2% at Max effort and 52.1% at Xhigh, against Opus 5.5's 54.4% and GPT-6 Sol's 49.3% as listed in the same table. Anthropic's own conclusion is measured: on several evaluations Sonnet 5.5 at Max effort performs comparably to Opus 5.5, but benchmark scores capture only one facet of a model's capabilities, and Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

The knowledge-work numbers show whose work it is really after. On GDPval-AA v2.1, a benchmark covering real-world work across 44 occupations, Sonnet 5.5 scores 1844 against Sonnet 5's 1449, Opus 5.5's 1846 and GPT-6 Sol's 1487 as listed in the same official table. On AA-Briefcase v1.1, a long-horizon knowledge work benchmark, Sonnet 5.5 scores 1811 against Sonnet 5's 1359 and Opus 5.5's 1822. In other words, the gap to the company's own flagship is down to two points. On computer use, OSWorld 2.1 at partial completion puts Sonnet 5.5 at 80.1% versus Sonnet 5's 57.0% and Opus 5.5's 81.8%. On the more visual task of chart recognition, Chartography with no tools puts Sonnet 5.5 at 61.6% against Sonnet 5's 15.6%, Opus 5.5's 64.4% and GPT-6 Sol's 53.6%.

The cost section is the real selling point here, and also the most easily misread. What Anthropic emphasizes is not a price cut but fewer tokens to do the same work: the price list stays at Sonnet 5 levels, while Anthropic says its testing showed up to 30% lower cost per task. The finer detail is in the accuracy-versus-cost charts. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task; on Terminal-Bench 4.0, Medium effort, the default in the Claude apps, far exceeds Sonnet 5's best score for less than a tenth of the cost; on CursorBench 4.0, Low effort exceeds Sonnet 5's best score; on AA-Briefcase, Medium effort bests Sonnet 5's best score for about one ninth of the cost. The usage guidance follows: Sonnet 5.5 complements Opus 5.5 best at lower effort settings, and at higher settings, where cost per task converges, teams should evaluate Opus 5.5 instead.

Two details on safety and migration are worth recording. Anthropic says that on its automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures, and that because its cybersecurity capabilities are comparable to Opus 5's, it is the first Sonnet model to launch with cyber safeguards and fallbacks like those developed for its most capable models, while its biology safeguards match Sonnet 5's. Both safeguard sets target a narrow set of high-risk requests, and routine software development and most life sciences work are unaffected. The engineering details are explicit too: the model ID is claude-sonnet-5-5, with a 1M-token context window and 128K maximum output, available on Amazon Web Services, Google Cloud and Microsoft Azure among other platforms, with zero data retention; if you previously ran Sonnet with thinking off, you need to switch to the between_tools setting before moving to Sonnet 5.5. Haiku 5.5 will join the family in the coming weeks, and Snowflake has announced Sonnet 5.5 on Cortex AI in public preview.

Anthropic also volunteered three details that work against its own headline, which deserves its own paragraph. First, on FrontierCode Sonnet 5.5 scores lower at Max effort than at Xhigh; the official explanation is that at Max effort it more often ran Claude Code's code-review skill, splitting the review across many subagents, and in two cases Cognition examined this led to a timeout or to edits beyond the task's scope, hence the lower score. Second, Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 that Anthropic found to have a bug which could degrade responses to requests using structured outputs; Anthropic expects the effect to be small, to understate Sonnet 5.5's performance, and says the bug has since been fixed. Third, Anthropic notes OpenAI recently fixed a bug that degraded image understanding in GPT-6 Sol, so official AA-Briefcase v1.1 and GDPval-AA v2.1 scores from Artificial Analysis, and Chartography scores from Surge AI, may not yet reflect the latest version of that model. Publishing your own weak spots and your rival's measurement lag is more useful than showing only the best-looking column.

One note on how to read a release like this. Placed back into late September's sequence, Anthropic's cadence is clear: the flagship Opus 5.5 on September 22 at 4 dollars per million input tokens and 20 dollars output, the mid-tier Sonnet 5.5 a week later on September 28 at half the price and within two points on the knowledge-work benchmark, with Haiku 5.5 still on the way. Reuters points out the backdrop: the company is expanding its product lineup ahead of a planned IPO, while its chief executive Dario Amodei earlier this month called on the global AI community to slow the pace of releasing new capabilities. The two facts are not contradictory. Pushing frontier capability down to mid-tier prices is itself a commercial choice that ties risk and adoption together. For buyers, the useful move is not arguing about which model is stronger but re-tiering your own tasks by difficulty and latency requirements, because a release where the price list holds steady while the effort-to-capability curve shifts changes your cost structure, not your habits.

🤔 Frequently Asked Questions

Did Sonnet 5.5 get more expensive?

No. Anthropic states pricing is the same as Sonnet 5: 2 dollars per million input tokens, 10 dollars per million output tokens and 0.20 dollars per million tokens for cache reads. The saving comes not from a price cut but from needing far fewer tokens to do the same work, with the official testing showing up to 30% lower cost per task.

Can it replace Opus 5.5?

Anthropic's answer is partly, depending on the scenario. On some evaluations Sonnet 5.5 at Max effort approaches Opus 5.5, for example 1844 versus 1846 on GDPval-AA v2.1, but the company states plainly that Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Its guidance adds that Sonnet 5.5 complements Opus 5.5 best at lower effort settings, and that at higher settings, where cost per task converges, teams should evaluate Opus 5.5.

How do you adopt it, and what migration notes apply?

The model ID is claude-sonnet-5-5, with a 1M-token context window and 128K maximum output, available on Amazon Web Services, Google Cloud and Microsoft Azure among other platforms, with zero data retention. One specific migration requirement: if you previously ran Sonnet with thinking off, you need to switch to the between_tools setting before moving to Sonnet 5.5, per Anthropic's migration guide.

Do the new cyber safeguards affect normal development?

Per Anthropic, the cyber safeguards and fallbacks match those developed for its most capable models, making this the first Sonnet model to launch with them, while biology safeguards are the same as Sonnet 5's. Both target a narrow set of high-risk requests, and routine software development and most life sciences work are unaffected. In other words, they intercept specific high-risk uses, not ordinary development tasks.

🛠️ Recommended Tools

  • AI Model RouterAnthropic's own guidance is routing logic: Sonnet at lower effort, Opus at higher effort. Move that decision out of business code into configuration, so a model refresh means editing a routing table rather than every call site.
  • AI Token CounterThe saving here comes from fewer tokens, not a lower rate, and cache reads are priced separately at 0.20 dollars per million. Measure the token distribution of real requests before you ship to see where that 30 percent actually comes from.
  • CSV to ExcelAnthropic lists polished documents, slides and spreadsheets among Sonnet 5.5's strong everyday scenarios, and that is also where enterprise demand concentrates. If your pipeline still moves data files by hand, automate that step before deciding whether to hand it to a model.

Summary

Per Anthropic's official launch page dated September 28 and Reuters: Claude Sonnet 5.5 is the second model in the Claude 5.5 family, running 30%+ faster than Sonnet 5 and costing up to 30% less for most work, with pricing held at Sonnet 5 levels (2 dollars per million input tokens, 10 dollars per million output tokens, 0.20 dollars per million tokens for cache reads), where the saving comes from needing fewer tokens. On benchmarks, Terminal-Bench 4.0 rises from 10.3% to 70.6% (Opus 5.5 at 66.4%); GDPval-AA v2.1 scores 1844 (versus Sonnet 5's 1449, Opus 5.5's 1846 and GPT-6 Sol's 1487); AA-Briefcase v1.1 scores 1811; OSWorld 2.1 scores 80.1%; Chartography scores 61.6%. It is the first Sonnet model to beat Pokemon Red working only from screenshots, and the first Sonnet model to launch with cyber safeguards and fallbacks. The model ID is claude-sonnet-5-5 with a 1M-token context window and 128K output, available on AWS, Google Cloud and Azure with zero data retention; Haiku 5.5 arrives in the coming weeks and Snowflake Cortex AI offers a public preview. Reuters adds that this is the company's second launch ahead of a planned IPO, with enterprise customers making up about 80% of its business. Anthropic also disclosed that FrontierCode scores lower at Max than Xhigh, that a pre-release deployment had a bug affecting structured outputs, and that a rival model's scores may not be updated. Every fact comes from the official and authoritative sources listed below, with no speculation added.

Sources: Anthropic: Introducing Claude Sonnet 5.5 (official)
Anthropic: Claude Sonnet 5.5 System Card (official)
Reuters: Anthropic rolls out second Claude 5.5 model as it builds toward IPO
Snowflake: Announcing Anthropic Claude Sonnet 5.5 on Snowflake Cortex AI
VentureBeat: Anthropic launches Claude Sonnet 5.5