Frontline Hotspot
Frontline Hotspot

Claude Sonnet 5.5 Ships: Coding Score Leaps From 10 to 70

Claude Sonnet 5.5 shipped September 28, 2026 (US Eastern, official basis): the second model in the Claude 5.5 family, positioned as a faster, cheaper complement to Opus 5.5, live day one on the Claude platform plus AWS Bedrock, Google Cloud Vertex and Microsoft Azure, with Claude Code integrated the same day. Headline numbers: Terminal-Bench 4.0 jumps from Sonnet 5's 10.3% to 70.6% (same benchmark, different generation - nearly sevenfold); CursorBench 4.0 at 55.5% (Opus 5.5: 57.8%); OSWorld 2.1 from 57.0% to 80.1%; GDPval-AA within 2 points of Opus 5.5; output speed up over 30% and per-task cost down up to 30% on most work. Pricing strategy: unit prices unchanged versus Sonnet 5 ($2/$10, cache read $0.20) - the discount hides in token efficiency. Security: the Sonnet line gets Opus/Fable-grade cyber protections for the first time, plus an anti-extraction classifier and auto-fallback for high-risk cyber requests. Also the first Sonnet to beat Pokemon Red from screenshots alone. Competitive context: Gemini 4 Argon landed two days later (controlled release, 1M output tokens) and GPT-6.1 Astra was delayed over safety; enterprise customers are about 80% of Anthropic's business ahead of a planned IPO (Reuters). Haiku 5.5 is teased for the coming weeks - no specs published, none invented here.

Published October 5, 20269 min read
<!-- claude-sonnet-5-5-release-hotspot | hotspot | Claude Sonnet 5.5 Ships: Coding Score Leaps From 10 to 70 -->

On September 28, 2026 (US Eastern time), Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, arriving a little over a week after Opus 5.5. Per the official positioning, Sonnet 5.5 is the faster, lower-cost complement to Opus 5.5, aimed at well-scoped everyday work: fixing bugs, writing documentation, and producing slides and spreadsheets. The model went live the same day on the Claude platform and on AWS Bedrock, Google Cloud Vertex, and Microsoft Azure, and Claude Code integrated it on day one.

The numbers are where this release gets interesting. Terminal-Bench 4.0, a coding benchmark, jumped from 10.3 percent for the previous Sonnet 5 to 70.6 percent for Sonnet 5.5. Official figures also claim output speed is up more than 30 percent and per-task cost on most workloads is down as much as 30 percent, all while the price list stays identical to Sonnet 5. A new model pulling nearly seven times higher on the same benchmark without raising unit prices by a cent deserves a closer read. This article breaks the launch down piece by piece: the official numbers, how to read the 10.3-to-70.6 jump correctly, the arithmetic behind the pricing strategy, the security and availability upgrades, and the competitive backdrop provided by Gemini 4 Argon, which landed two days later. All figures come from official statements or clearly marked media-cited sources; anything unpublished is not filled in with guesses.

What Shipped: The Official Numbers, One by One

The official framing for Sonnet 5.5 is deliberately modest: it is the faster, cheaper complement to Opus 5.5, built for well-scoped daily tasks, bug fixes, documentation, slides, and spreadsheets. In plain terms, Opus 5.5 handles the long and complicated jobs, while Sonnet 5.5 handles the work you do every day. The real substance sits in the measurable figures, so here they are in order.

First, speed. Official statements put output speed at more than 30 percent faster than Sonnet 5. For high-frequency coding use, 30 percent faster output compresses waiting time directly, and dozens of small calls per engineer per day add up to a noticeable difference.

Second, cost. Anthropic says most workloads see per-task costs drop by as much as 30 percent. The source of the saving is not a price cut but fewer tokens: the same task simply consumes less output. This point and the pricing strategy are two sides of the same coin, covered in detail below.

Third, the coding benchmarks. Sonnet 5.5 scores 70.6 percent on Terminal-Bench 4.0, where the previous Sonnet 5 managed only 10.3 percent. That near-sevenfold jump on the same benchmark is the headline of this launch. On CursorBench 4.0 it scores 55.5 percent, close to Opus 5.5's 57.8 percent. In other words, on coding, the volume tier has nearly caught up to where the flagship stood one generation ago.

Fourth, the other benchmarks. OSWorld 2.1 climbs from 57.0 percent to 80.1 percent, a gain of 23 points. Chartography rises from 15.6 percent to 61.6 percent, close to four times higher. On GDPval-AA it trails Opus 5.5 by only 2 points. Taken together, the picture is one of a gap to the flagship that has narrowed across the board, nearly closing on some dimensions.

Fifth, third-party signals. Media reports cite Artificial Analysis intelligence scores placing Sonnet 5.5 above GPT-6 Astra and Claude Fable 5.1. That figure deserves a clear label: it is a media-cited third-party number, not an independent test by this site. For a fuller picture of the GPT-6 Astra side of the ledger, see our earlier review of GPT-6 Astra.

Sixth, a fun footnote. Both official material and TechNews note that Sonnet 5.5 is the first Sonnet model to complete Pokemon Red using screenshots alone. It will not factor into any purchasing decision, but it is a vivid side profile of long-horizon planning: sustaining hundreds of consecutive decisions without derailing bodes well for agent-style applications.

10.3 to 70.6: Same Benchmark, Different Generations

Start with the reading that trips people up most often. The 10.3 percent and the 70.6 percent come from the same benchmark, Terminal-Bench 4.0, applied to the previous Sonnet 5 and the new Sonnet 5.5 respectively. Same exam, two generations of students. It is not a collage of scores from different tests stitched together for dramatic effect.

Why stress this? Because the most common numbers game of launch season is precisely the benchmark swap: the new model runs the new benchmark, the old model keeps its old score, and the flattering figure goes on the front page. Here Terminal-Bench 4.0 tested both generations on the same paper, so the jump is genuinely comparable. A near-sevenfold gap does not describe incremental progress; it describes a generational shift in terminal-side coding, from occasional success to majority completion. Where the previous model might succeed once in ten attempts, the new one lands seven out of ten. Those are two entirely different definitions of usable.

The opposite misreading needs guarding too. Comparison tables circulating online increasingly place scores from different benchmark versions in a single column. A simple defense: Terminal-Bench 3.0 and Terminal-Bench 4.0 are different benchmark versions with different difficulty and test sets, and their scores cannot be compared directly. Check the version number before you read the number. Also note that every benchmark figure above is a vendor-reported figure for now; independent verification will take time, and this site does not arbitrate without running its own tests.

Pricing: Same Unit Price, Savings Hidden in Fewer Tokens

The price list is the most intriguing part of this launch. Sonnet 5.5 matches Sonnet 5 exactly: 2 dollars per million input tokens, 10 dollars per million output tokens, and 0.20 dollars per million cached-read tokens. Not a cent off the unit price, yet the official line simultaneously claims per-task costs on most workloads drop by up to 30 percent. Put the two sentences together and one strategy emerges: the discount does not come from the price tag, it comes from spending fewer tokens.

This route differs fundamentally from a straightforward price cut. A price cut charges less for the same amount of work. Token efficiency turns the same job into less work: fewer detours, less filler, fewer wasted retries, and the bill falls on its own. For users, efficiency-based savings are often the better deal, because they arrive automatically with the model upgrade; no renegotiation, no plan switch, no contract change. The trade-off is that the saving depends on the task type. The official phrasing is "most workloads" up to 30 percent, not a guaranteed 30 percent on everything, so leave room for variance when you budget.

For reference, Opus 5.5 is priced at 4 dollars per million input tokens and 20 dollars per million output tokens (official figures), positioned for complex long tasks. Sonnet 5.5 takes on everyday engineering work at less than half that, and the tiering of the product line is clear. Our review at the end of September, The Price Shuffle in Coding Models, already ran the numbers on Sonnet 5.5 against GPT-6.1 Sol. One reminder: all prices here are snapshots at publication time and defer to the official pages.

Security and Availability: Sonnet Gets Opus-Grade Protections

Two strands of this upgrade are easy to overlook but carry real weight.

The first is security. Official statements say Sonnet 5.5 is the first Sonnet to receive cybersecurity protections at the same tier as Opus and Fable, with cyber-offense and defense capability on par with Opus 5. Biological protections stay at the Sonnet 5 standard (CB-1). A new safety classifier guards against reasoning extraction, and high-risk cybersecurity requests can automatically fall back to Sonnet 5. Pushing flagship-grade protections down to the volume tier signals a clear judgment: models like this will run in real production at scale, and protection levels must track usage volume rather than model tier. The model with the largest attack surface is the one that runs the most.

The second is availability. The model ID is claude-sonnet-5-5, with a knowledge cutoff of June 2026, matching Opus 5.5. It went live on launch day on the Claude platform and on AWS Bedrock, Google Cloud Vertex, and Microsoft Azure, with Claude Code integration the same day and a zero data retention option on offer. The significance of the three-cloud rollout is that enterprises can plug in through cloud agreements they already have, with no new account systems, keeping existing procurement and compliance processes intact. For the hands-on details on how to call the model and where the entry points are on each cloud, see our companion guide, How to Call Claude Sonnet 5.5: Three Clouds, One Price List.

One more teaser worth recording: Anthropic says Claude Haiku 5.5 arrives within weeks, positioned as the high-throughput, low-cost tier. Its specifications are unpublished, so we will not invent them.

Competitive Context: Gemini 4 Argon Lands Days Later

Two days after Sonnet 5.5, on September 30 (US Eastern), Google released Gemini 4 Argon, pushing the single-response output limit from 64,000 tokens to one million. The two product bets point in opposite directions. Argon wagers on finishing long tasks in one breath; Sonnet 5.5 wagers on making each task faster and cheaper. The market received two opposite answers in the same week, and that tension is the story of the quarter: long-context work cares about output ceilings, everyday engineering cares about value for money, and the two roads have not yet collided head-on.

Argon's availability status needs an honest note: it is a controlled, staged release, currently limited to the Fairwind program of more than 650 cybersecurity partners and Google's internal use. The next stage targets paid API customers and AI Ultra subscribers, with no announced date. Ordinary users and developers cannot buy it yet. The independent picture deserves equal billing: Artificial Analysis testing shows Argon slightly behind its official numbers, with its Agent Arena coding result trailing Opus 4.8 and GPT-5.6-Sol. Official and independent calibers coexist, and firmer conclusions should wait for general availability. On pricing, Argon's introductory rate is 2 dollars per million input tokens and 10 dollars per million output tokens, rising later to 4 and 20 dollars (officially announced plan).

One piece of more distant background: Reuters reported on September 28 that GPT-6.1 Astra's release was delayed for safety reasons. The sway between shipping fast and shipping carefully among the major labs defined September, and our earlier piece on OpenAI's safety slowdown reads well alongside it. For how several models stack up on long output specifically, see our companion comparison, One Million Output Tokens, which concludes by need rather than by untested arbitration.

The Takeaway

Close with the capital backdrop. Reuters reports that enterprise customers account for roughly 80 percent of Anthropic's business, and media coverage describes the company as in the run-up to an IPO. Against that background, the play behind Sonnet 5.5 comes into focus: hold the unit price, lift efficiency, and push enterprise-grade protections down to the volume tier. Every move points toward making enterprises want to renew for the long term, and steady call volume is what a pre-IPO income statement needs.

Our read of this generation is that its bet is not benchmark supremacy but better engineering delivery at the same price. The 70.6 percent jump next to an unchanged price list is the best footnote to that bet. How the scores hold up will depend on independent testing and real production feedback. What can be said now is that the value-for-money bar for everyday coding tasks has moved up a notch.

Two adjacent stories are worth a look. NVIDIA just open-sourced a runtime for locking down AI agents along with a chip-level watchdog; see NVIDIA Locks Down AI Agents. And if you want the calling path mapped out before you start, go straight to the three-cloud API guide above.

Join the Discussion

Have you started using Sonnet 5.5? Does the 70.6 percent coding score hold up on your real tasks, and did your per-task cost actually drop by 30 percent? Tell us about your hands-on results in the comments.

Sources

  • Anthropic official site and transparency page (primary sources): launch information, benchmark scores, pricing, and security mechanisms
  • Reuters, September 28, 2026: launch coverage, GPT-6.1 Astra delay, enterprise revenue share
  • TechNews (Taiwan), September 29, 2026: three-cloud availability, the Pokemon Red claim
  • Synced (Machine Heart): cross-check of launch information
  • Artificial Analysis: independent intelligence scores and Argon independent testing (both media-cited figures)
  • Note: all benchmark scores and prices are snapshots at publication time and defer to the official Anthropic pages; Argon's price increase and rollout plan are officially announced projections.

This article is AI-assisted and human-edited. Last updated: 2026-10-05

Related

Frontline Hotspot

Anthropic's $2 Trillion IPO Run Starts With a $30 Trillion Pitch to Wall Street

Per the Wall Street Journal on August 25, Anthropic's IPO filing presents a total addressable market of over $30 trillion - surpassing SpaceX's $28.5 trillion to become the largest market narrative in business history. The math does not start from software or hardware sales: it tallies the total economic value and labor cost of all future work replaceable by AI models. The listing targets Fall 2026, a raise of up to $100 billion, and a valuation anchored near $2 trillion - a 2.2x jump over the $900 billion private valuation from May. The confidence: $11.6 billion Q2 2026 revenue and first positive adjusted operating profit; the risks: constrained US data center construction, overseas low-price model competition, and shrinking secondary-market tolerance. September's formal prospectus is the next hard milestone.

Aug 26, 20266 min read