Artificial intelligence
Anthropic's Mid-Tier Model Beats Its Six-Day-Old Flagship at Agentic Coding, for Half the Price. The Ladder Has Stopped Meaning Anything
Claude Sonnet 5.5 shipped on September 28 at unchanged pricing and, on Anthropic's own benchmarks, outscores the Opus 5.5 flagship on agentic coding while costing half as much. The quieter news is a classifier built to stop anyone extracting the model's reasoning.
MAI
On September 28 Anthropic released Claude Sonnet 5.5 at exactly the price Sonnet 5 had been selling for: $2 per million input tokens and $10 per million output. On the company's own benchmark table, the new mid-tier model scores 70.6% on Terminal-Bench 4.0. Claude Opus 5.5, the flagship released six days earlier and priced at $4 and $20, scores 66.4%.
That single comparison is more interesting than anything in the launch copy. Anthropic sells a three-tier ladder — Haiku, Sonnet, Opus — and the ladder is supposed to trade capability against cost. On the agentic coding benchmark Anthropic chose to lead with, the middle rung now sits above the top one at half the price.
What the numbers say
Every figure below comes from Anthropic. No independent evaluation of Sonnet 5.5 existed at the time of writing.
| Benchmark | Sonnet 5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 10.3% | 70.6% | 66.4% |
| CursorBench 4.0 | — | 55.5% | 57.8% |
| GDPval-AA v2.1 (knowledge work) | 1449 | 1844 | 1846 |
| AA-Briefcase | — | 1811 | 1822 |
| OSWorld 2.1 (computer use) | 57% | 80.1% | 81.8% |
| Model | Input / MTok | Output / MTok |
|---|---|---|
| Claude Opus 5.5 | $4 | $20 |
| Claude Sonnet 5.5 | $2 | $10 |
| Claude Sonnet 5 | $2 | $10 |
Read the GDPval-AA row carefully: 1844 against 1846 is not a gap, it is measurement noise. On four of the five benchmarks Sonnet 5.5 lands within two or three points of a model that costs twice as much, and on the fifth it wins.
The Terminal-Bench row deserves scepticism rather than applause. A jump from 10.3% to 70.6% inside one point release is not a capability curve; it is a model that was not built for that harness being replaced by one that was. Sonnet 5's score tells you how badly it handled a particular agentic scaffold, not how incapable it was. Treat the 70.6% as evidence that Anthropic optimised for the thing it is now measuring.
The efficiency claims sit on the same footing. Anthropic says Sonnet 5.5 generates output about 30% faster than Sonnet 5 and cuts cost per task by up to 30% through shorter outputs and fewer tool calls, and it published customer figures to match — Box reporting 2.4x faster runs with 12% fewer tokens, Slack 14% fewer output tokens, Lovable a third fewer tool calls, Base44 completing tasks in 3.6 iterations against 7.7 for Opus 5. These are numbers supplied to a vendor by design partners ahead of a launch. They point in a consistent direction, which is worth something, but none of them is an independent measurement.
What the Opus tier is now for
If the mid-tier model matches the flagship on general knowledge work and beats it on agentic coding, the obvious question is what the extra $2 and $10 buy. On Anthropic's own published pricing the answer is narrowing: Opus 5.5 retains a cheaper cache-read multiplier and a "fast mode" at $8 and $40, and it stays marginally ahead on CursorBench, AA-Briefcase and OSWorld. That is a real but thin set of reasons.
The more honest reading is that the frontier tier has become a place to put capabilities before they are cheap enough to ship at volume, and the interval between the two is now measured in days rather than quarters. Six days separated Opus 5.5 from a model that undercuts it. Anyone building on the Opus tier for cost-insensitive reasons should be re-running their own evaluations rather than trusting the ladder.
The classifier nobody led with
Buried in the safety notes is the launch's second story. Sonnet 5.5 is the first Sonnet model to carry Anthropic's cybersecurity safeguards, routing high-risk requests back to Sonnet 5. It also ships classifiers that block extraction of the model's reasoning.
That last one is not a safety feature in the usual sense. Reasoning traces are the raw material of distillation — the technique by which a cheaper model is trained on an expensive model's outputs, and the technique that produced several of the open-weight competitors now pressing Anthropic on price. Anthropic is treating its own chain of thought as an asset to be defended, and has now said so in a product announcement.
It lands in a week when OpenAI pulled GPT-6.1 Astra over safety-test regressions. One lab held a frontier model back; the other shipped a cheaper one and locked the doors on its reasoning. Both are responses to the same pressure, which is that capability is no longer the scarce thing. Delivered cost is.
Sources: Introducing Claude Sonnet 5.5 — Anthropic · Claude Platform pricing · Anthropic launches Claude Sonnet 5.5 with 30% cost reduction per task — VentureBeat · Anthropic Releases Claude Sonnet 5.5 at Unchanged Sonnet 5 Pricing — Unite.AI · Anthropic releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 — MarkTechPost