Anthropic released Claude Sonnet 5.5 on September 28 and left Sonnet 5's pricing unchanged: $2 per million input tokens, $10 per million output. The independent Artificial Analysis index ranks the new model second in the world, two points behind Opus 5.5 and three ahead of GPT-6 Astra. The same index also shows something the press release does not: at maximum effort, one task costs about $7.60 with Sonnet 5.5, more than with Opus 5.5. The price per token stayed put. The bill did not.
What you get
The model ships as claude-sonnet-5-5, with a 1M-token context window, up to 128K tokens of output and a June 2026 knowledge cutoff. Adaptive reasoning is on by default and there are five effort levels, from low to max. The weights remain closed. Sonnet 5.5 runs in the Claude apps, on the Claude Platform, and in AWS, Google Cloud and Microsoft Azure. Caching costs $0.20 per million tokens to read and $2.50 to write.
Anthropic says the model generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task, despite the identical per-token rate. The saving is supposed to come from needing fewer tokens to solve the same problem. The company cites customers: Balyasny Asset Management measured roughly 121K tokens per answer against 497K on Sonnet 5, Box recorded 2.4x faster results with 12% fewer tokens, Slack saw about 14% fewer output tokens, and Atlassian says its Rovo agents run up to 30% faster.
Results: Anthropic versus outside measurement
On Anthropic's own tests the jump is enormous. Terminal-Bench 4.0, an agentic terminal coding test, gives Sonnet 5.5 70.6% against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. CursorBench 4.0 shows 55.5% against 34.1%, FrontierCode 1.1 at max effort 46.2% against 42.4%, and OSWorld 2.1 for computer use 80.1%. Anthropic also says it is the first Sonnet to beat Pokémon Red from screenshots alone.
Artificial Analysis measures the same things differently and gets smaller numbers with the same ordering. Terminal-Bench 4.0 at max effort: 64% for Sonnet 5.5 against 60% for Opus 5.5 and 60% for GPT-6 Astra. The lead over the company's own flagship is real here, just smaller than the announcement suggests. The overall index score is 56, up 18 points from Sonnet 5's 38.
The table shows where the "smaller model" label stops applying. Sonnet 5.5 wins where autonomous work with tools decides the score, and loses to Opus on factual knowledge: 54% accuracy against 66%. It also hallucinates less, because it declines more often than it guesses, which for agent users is often worth more than the raw score. Artificial Analysis cautions that it tested a pre-release build with a structured-outputs bug fixed in the public release, and says it will re-run the evaluation.
The bill per task
The most interesting number from the independent measurement is not about quality. At max effort, Sonnet 5.5 used an average of about 193K output tokens per index task, the most of any model Artificial Analysis has measured. That is roughly 60% more than Opus 5.5 and about seven times GPT-6 Astra at max. At the same per-token rate, a task costs about $7.60, roughly half again what it cost with Sonnet 5.
This does not contradict Anthropic's claim of 30% lower cost per task. The company compares against Sonnet 5 on its customers' own workloads and at lower effort levels, where the model does save tokens. A ComputingForGeeks test shows how wide the spread is: on the same deployment, max effort generated about 22 times more tokens than medium, and medium passed every test in 36 seconds. In the same test, a debugging task in Claude Code took Sonnet 5.5 three tool calls against 12 to 13 for its predecessor.
In our view that is the real headline: for models with a "thinking" mode, the price per token has stopped being the price of the product. The actual cost dial is the effort level, and its default setting shapes the bill more than the price list. Anyone who switches Sonnet 5.5 to max out of habit from the previous model will pay more for Sonnet than for Opus.
Filters, refusals and the fallback to Sonnet 5
Sonnet 5.5 is the first Sonnet with cyber safeguards at the Opus 5 level. The system card reports 99.43% recall on cyber harm coverage and a "rewind" attack success rate falling from 57.2% on Sonnet 5 to 21.0%. The price of that is written plainly in the documentation: users should expect more refusals than with Sonnet 5, including on legitimate security work.
Requests flagged as cyber, frontier model development (for example kernels for ML accelerators) or attempts to extract reasoning are automatically rerouted to Sonnet 5, and the user sees a notice about the model switch. Biology requests are blocked with no fallback. In the apps switching is on by default and can be turned off in Settings so the chat simply stops. In the API automatic fallback is not active by default and must be enabled in configuration. The Cyber Verification Program, which gives defenders access to fuller capabilities, does not cover Sonnet 5.5 at launch. Anthropic promises to extend it soon.
We know this pattern from Fable 5.1 and Mythos 5.1, where two-tier access became a permanent system. Sonnet 5.5 carries it into the mid tier, which had been free of such filters. For pentesters and security researchers it means the cheapest model in the line stops being a good tool for their work until the verification program expands.
API migration
According to the ComputingForGeeks test, the release introduces breaking changes: forced tool use is removed, manual thinking budgets are rejected, and the temperature parameter is being deprecated. Thinking cannot be turned off, only reduced to reasoning between tool calls. Anyone with a production integration built on Sonnet 5 should test it before switching the alias. Moving to Sonnet 5.5 requires changes on the client side.
For the comparison with the rest of the top tier, see our piece on the Opus 5 launch, in which Anthropic first sold near-frontier intelligence for a fraction of the price.
Sonnet 5.5 is one of the strongest models for agentic work in the terminal and in tools, and the first Sonnet to compete with its own company's flagship. That is an achievement, and the independent measurement confirms it better than the press release, because it shows a lead over Opus 5.5 on Terminal-Bench, though a smaller one than claimed.
We would still warn against reading the price list the way you did with earlier generations. Leave effort at max and you get the most expensive bill in this class. Start at medium, check on your own tasks how many tokens the model really burns, and raise effort only where a test shows it changes something. The default configuration is the biggest line item on the invoice.
The largest hidden cost is refusals. For an ordinary developer the fallback to Sonnet 5 will rarely be visible, but for anyone working near security, Sonnet 5.5 without the Cyber Verification Program can be a worse choice than its predecessor. If Anthropic expands the program quickly, that downside disappears. Until then we recommend Sonnet 5.5 for code, automation and agentic work, and advise against it for offensive and research cyber.





Comments
Discussion
Join the conversation around this story.
Join the discussion
Sign in to comment and reply to other readers.
Sign inNo comments yet
Start the discussion first.