
Anthropic just dropped Claude Haiku 5.5 on October 7, 2026 (UTC), and the numbers are hard to ignore. Their smallest model now beats OpenAI's GPT-6 Luna on nearly every shared benchmark — at roughly one-tenth the price. This isn't a minor tweak. It's a deliberate shot across the bow at the budget-tier model market, and it changes what "good enough" means for high-volume AI workloads.
What's new
Haiku 5.5 (model ID: claude-haiku-5-5) is Anthropic's cheapest, fastest, and most capable small model to date. It's built for the boring-but-enormous part of the AI stack: summaries, compactions, database queries, classification, subagent work in coding pipelines, live customer support, and browser automation.
The headline improvements fall into three buckets:
- Performance: It clears GPT-6 Luna on GDPval-AA, AA-Briefcase, OSWorld, Terminal-Bench, FrontierCode, and Chartography. On OSWorld 2.1 (computer use), Haiku 5.5 scores 72.4% versus Luna's 48.9% — a 48% relative jump. On Terminal-Bench 4.0 (agentic coding), it hits 39.2% versus Luna's 16.4%, while Haiku 4.5 sat at 0.0%.
- Price: Input tokens drop from $1.00 to $0.10 per million for prompts under 100k tokens — a 90% cut. Output falls from $5.00 to $0.50. For prompts over 100k, the discount is 50%. Anthropic says the average workload now costs about 75% less than Haiku 4.5.
- Effort control: It's the first Haiku-class model with an adjustable effort setting (Low through Max), letting users trade intelligence for cost on the fly.
Alongside the launch, Anthropic halved Sonnet 5.5's cache-read price from $0.20 to $0.10 per million tokens, which reduces most agentic workloads by roughly 20%. Max and Team subscribers also get monthly API credits — $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team plans.
Benchmark comparison
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 (coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (coding) | 46.4% | — | 42.4% | 52.1% |
| Chartography (visual) | 46.4% | 6.4% | 29.1% | 61.6% |
| HLE no tools (reasoning) | 45.9% | 10.2% | — | 56.9% |
Source: Anthropic official announcement, October 7, 2026 (UTC). "—" indicates the competitor did not publish a comparable score.
Pricing comparison (per 1M tokens, prompts ≤100k)
| Tier | Cache read | Cache write | Input | Output |
|---|---|---|---|---|
| Haiku 5.5 | $0.01 | $0.125 | $0.10 | $0.50 |
| Haiku 4.5 | $0.10 | $1.25 | $1.00 | $5.00 |
| Sonnet 5.5 | $0.10 | $2.50 | $2.00 | $10.00 |
Source: Anthropic pricing page, October 7, 2026 (UTC).
Why it matters
The small-model tier is where the volume lives. Most production AI systems don't send every request to a flagship model — they route summaries, extractions, and routing decisions to cheaper models. Haiku 5.5's combination of beating GPT-6 Luna on benchmarks while costing 90% less on input puts real pressure on OpenAI's pricing in exactly the segment where OpenAI has the most to lose.
Look at the OSWorld number. A 72.4% score on computer-use tasks from a $0.10-input model was not realistic six months ago. That capability unlocks browser automation and desktop agent workflows at a price point where unit economics finally work. Cognition already integrated Haiku 5.5 as a sidekick in Devin Fusion, holding a 66.2 FrontierCode score while cutting cost and latency. Asana reported 30% lower latency and 2.5x faster inference per agent turn. HubSpot saw 92.8% on its CRM eval suite — the best score they've recorded.
The strategic read is that Anthropic is building a full-stack price ladder. Opus 5.5 owns the high-intelligence ceiling, Sonnet 5.5 covers the middle, and now Haiku 5.5 aggressively defends the volume floor. The cache-read cut on Sonnet 5.5 is not a coincidence — it makes the two-model routing pattern (Haiku for subagents, Sonnet for the main task) materially cheaper, which locks customers deeper into the Claude Platform.
The catch
Haiku 5.5 is not a flagship replacement. On Terminal-Bench 4.0, Sonnet 5.5 still scores 70.6% versus Haiku's 39.2%. On GDPval-AA, Sonnet leads 1840 to 1620. For complex, multi-step agentic coding, the bigger model remains the right call. Anthropic's own framing is explicit: Haiku 5.5 is for "narrowly scoped tasks that might otherwise have been cost-prohibitive."
The 100k-token pricing cliff is worth watching. Prompts under 100k get the 90% discount; anything over jumps to $0.50 input and $2.50 output. Anthropic notes that 90% of Haiku 4.5 requests fell under 100k, so most users won't hit it — but teams doing long-context compaction or RAG over large documents need to recalculate.
On safety, Haiku 5.5's cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat looser than Sonnet 5.5's. They permit a wider range of defensive tasks while still blocking penetration testing. The system card reports a 38% capability rate on ExploitBench's plain arm and 49% with AutoNudge, with full arbitrary code execution 1.0% of the time.
What to watch
- OpenAI's response: GPT-6 Luna now loses on both performance and price in the small-model tier. A Luna price cut or a GPT-6.1 Luna refresh before the end of Q4 2026 would be the tell.
- Adoption metrics: If Haiku 5.5's share of Anthropic's API traffic crosses 40% within two months, it validates the price-war thesis. Watch Anthropic's next developer update for token-volume disclosures.
- Open-weight pressure: Mistral Large 4 arrives in open weights on October 27, 2026 (UTC). Haiku 5.5 is closed but cheap — the question is whether budget buyers prefer predictable API pricing or self-hosted flexibility.
- Sonnet 5.5 cache economics: The 50% cache-read cut makes long-running agents significantly cheaper. If competitors don't match, Anthropic could pull agentic-workload share from both OpenAI and Google in Q4.
Haiku 5.5 is available now on the Claude Platform, AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure.
No comments yet