
Anthropic released Claude Sonnet 5.5 on September 28, 2026 (UTC), and the numbers are hard to ignore. It's a mid-tier model that posts flagship-level scores — in some benchmarks it even beats Opus 5.5, Anthropic's own top model — while costing half as much. The timing isn't subtle either: it landed the day before OpenAI's DevDay.
The headline numbers
Sonnet 5.5 isn't a gentle refinement. On Terminal-Bench 4.0, an agentic coding test, it scores 70.6%. Sonnet 5 scored 10.3%. That's not an incremental bump — it's a nearly sevenfold jump, and it edges past Opus 5.5's 66.4% at Xhigh effort.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 (Xhigh) | 52.1% | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | — |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
| Humanity's Last Exam (tools) | 64.5% | 54.9% | 67.7% | — |
Source: Anthropic official announcement, September 28, 2026 (UTC). GPT-6 Sol figures are Anthropic's own runs; Terminal-Bench and CursorBench don't have public GPT-6 Sol scores, so Anthropic reports GPT-5.6 Sol there instead.
On the Artificial Analysis Intelligence Index, Sonnet 5.5 scores 56 — two points behind Opus 5.5 at 58, but ahead of GPT-6 Astra at 53 and GPT-6 Sol at 48. That means Anthropic now holds the top two spots on a widely cited independent leaderboard.
Cost and speed
The price didn't move: $2 per million input tokens, $10 per million output, $0.20 for cache reads. Same as Sonnet 5. But Sonnet 5.5 needs far fewer tokens to finish the same work — Anthropic says up to 30% less per task — and generates output 30%+ faster. It's the first Sonnet model to ship with the same cyber safeguards and fallbacks that Opus-class models get, because its offensive security capabilities are now comparable to Opus 5's.
It's also the first Sonnet to beat Pokémon Red working only from screenshots, which sounds like a gimmick but is actually a decent proxy for long-horizon visual planning.
Why it matters
The real story here isn't "mid-tier model gets better." It's that the gap between Anthropic's second-best model and OpenAI's flagship has collapsed. Sonnet 5.5 matches or beats GPT-6 Sol on every benchmark where both are measured, and on GDPval-AA — OpenAI's own knowledge-work benchmark — it scores 1844 versus GPT-6 Sol's 1487. That's a 24% gap in Anthropic's favor, on a test OpenAI designed.
The pricing is the sharper weapon. Sonnet 5.5 costs half what Opus 5.5 does ($2/$10 vs $4/$20 per million tokens) while delivering 95-98% of the performance on most tasks. For enterprise buyers running high-volume agent workloads, that's not a rounding error — it's a reason to switch. Early testers at Slack, Zendesk, and Balyasny Asset Management all reported fewer tokens, faster resolution, and better quality than Sonnet 5, with no prompt changes needed.
The DevDay timing is the cherry on top. OpenAI is expected to showcase new models and agent products at its event starting at 10:00 (UTC-7) on September 29, 2026. Anthropic just dropped a model that says, in effect: whatever you announce, we already have a cheaper alternative that scores higher on your own benchmark.
Critical thinking
A few caveats. Most of these benchmarks are Anthropic's own runs, and the company has an obvious incentive to pick tests where it looks good. The Artificial Analysis Intelligence Index is independent, but it's one composite score, not ground truth. And "beats GPT-6 Sol" is not the same as "beats GPT-6 Astra" — Astra remains the more capable model on open-ended reasoning, even if Sonnet 5.5 closes the gap on structured tasks.
There's also the token-efficiency claim. Anthropic says Sonnet 5.5 uses fewer output tokens per task, but Artificial Analysis notes that at Max effort it burns ~193k output tokens per Intelligence Index task — the highest they've seen. The efficiency gain may be real at default settings but evaporate when users crank effort to maximum.
What to watch
- How OpenAI responds at DevDay — expect pricing pressure and a direct benchmark rebuttal
- Whether Sonnet 5.5's Terminal-Bench lead holds up under independent testing
- Haiku 5.5, due in the coming weeks, which could compress the cost-performance ladder even further
- Enterprise migration data: if Slack and Zendesk roll Sonnet 5.5 into production broadly, that's a real revenue signal
Sources: Anthropic official announcement (anthropic.com/claude-sonnet-5-5), Artificial Analysis Intelligence Index, The Decoder, Digital Trends.
No comments yet