Benchmarks · updated October 8, 2026
Haiku 5.5 benchmark results
Anthropic launched Claude Haiku 5.5 on October 7, 2026 with eight headline benchmark scores. Here is every one of them next to Haiku 4.5, Sonnet 5.5 and OpenAI's GPT-6 Luna: what each test measures, how big the jump is, and what independent testers found.
- OSWorld 2.1, computer use. Haiku 4.5: 15.7%
- 72.4%
- Terminal-Bench 4.0, agentic coding. Haiku 4.5: 0.0%
- 39.2%
- Humanity's Last Exam with tools. Haiku 4.5: 18.7%
- 57.4%
- GDPval-AA Elo, knowledge work. Haiku 4.5: 735
- 1620
All results at a glance
Anthropic's published table, unchanged. Higher is better on every row. GDPval-AA and AA-Briefcase are Elo ratings; the rest are percentages.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1Knowledge work · Elo rating | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1Knowledge work · Elo rating | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1Computer use · offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last ExamReasoning · no tools | 45.9% | 10.2% | Not reported | 56.9% |
| Humanity's Last ExamReasoning · with tools | 57.4% | 18.7% | Not reported | 64.5% |
| Terminal-Bench 4.0Agentic coding · pass@1 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1Agentic coding · Main | 46.4% | Not reported | 42.4% | 52.1%xhigh effort |
| ChartographyVisual reasoning · no tools | 46.4% | 6.4% | 29.1% | 61.6% |
Source: Anthropic, “Introducing Claude Haiku 5.5”, October 7, 2026. A dash means not reported. Sonnet 5.5's FrontierCode score was run at xhigh effort. Evaluation details are in the Haiku 5.5 system card.
Haiku 5.5 vs Haiku 4.5: a generational jump
Haiku 4.5 was a fast chat model; Haiku 5.5 is a working agent. The biggest gains are on tests where the model has to act, not just answer.
On Terminal-Bench 4.0, which has the model complete multi-step professional tasks in a command line, Haiku 4.5 scored 0.0% in Anthropic's run and Haiku 5.5 scores 39.2%. On OSWorld 2.1, where it operates a real computer through long tasks, it climbs from 15.7% to 72.4%. And it gets there at up to 90% lower cost.
Haiku 5.5 vs GPT-6 Luna
GPT-6 Luna is OpenAI's small model and the rival Anthropic chose to compare against. Haiku 5.5 leads on all six benchmarks where both have a score.
Independent testing is closer. Artificial Analysis gives Haiku 5.5 a 43 on its Intelligence Index at max effort, ahead of GPT-6 Luna's 38. At high effort the two tie at 38, and at max effort Haiku 5.5 used about three times as many output tokens as Luna. Artificial Analysis also measured a lower hallucination rate for Haiku 5.5 (40% vs 77%) and higher knowledge accuracy for Luna (44% vs 36%).
How close is Haiku 5.5 to Sonnet 5.5?
Sonnet 5.5 stays ahead on every benchmark. The real question is how much of Sonnet's score you get for 1/20 of its per-token price.
On reasoning, computer-use and graded coding tests, Haiku 5.5 reaches 81–89% of Sonnet 5.5's score. On long agentic coding the gap is real: 39.2% vs 70.6% on Terminal-Bench 4.0. That matches Anthropic's own advice: Sonnet 5.5 or Opus 5.5 for complex agentic coding, and Haiku 5.5 for high-volume work and as their sub-agent.
What about SWE-bench?
Anthropic's launch post does not include SWE-bench; its coding benchmarks are Terminal-Bench 4.0 and FrontierCode. Launch coverage citing the Haiku 5.5 system card reports 64.8% on SWE-bench Pro and 83.7% on SWE-bench Multilingual.
You may also see Haiku 4.5 quoted at 73.3%. That score is on SWE-bench Verified, a different test set, so it can't be compared with the numbers above.
Effort changes the score
Haiku 5.5 is the first Haiku with effort levels: low, medium (the default), high, xhigh and max. Anthropic charts each level's accuracy against its cost on OSWorld, GDPval-AA and Humanity's Last Exam. Scores climb with effort, and so does token use.
Artificial Analysis saw the same curve: Haiku 5.5 scores 38 on its Intelligence Index at high effort and 43 at max, and the step from xhigh to max adds two points for about 1.8× the tokens. For everyday chat, low or medium effort is the sweet spot; save the top levels for hard problems. On Haiku55, the Fast, Smart and Deep modes map to low, medium and high effort.
What early testers measured
Results that companies reported from their own evaluations at launch:
HubSpot
92.8% on its CRM simulation suite, averaged over three runs: the best score HubSpot has seen on it.
AlphaSense
0.84 versus 0.76 for Haiku 4.5 across 400 queries on its Ask in Document feature.
Box
11 points higher than Haiku 4.5 at about half the latency.
Asana
Over 30% lower task-completion latency and up to 2.5× faster inference per agent turn.
What each benchmark measures
- GDPval-AA v2.1Knowledge work
- Real professional work drawn from 44 occupations. Models work as agents, and blind head-to-head grading ranks them on an Elo scale. Run by Artificial Analysis.
- AA-Briefcase v1.1Knowledge work
- Artificial Analysis's private test of long-horizon business work, where agents produce deliverables such as spreadsheets, presentations and memos. Also scored as Elo.
- OSWorld 2.1Computer use
- The model operates a real computer to finish long, multi-step tasks across desktop apps. Anthropic reports Haiku 5.5 on the offline subset.
- Humanity's Last ExamReasoning
- Humanity's Last Exam: expert-level questions across dozens of academic fields, written to be hard for frontier models. Reported both with and without tools.
- Terminal-Bench 4.0Agentic coding
- Multi-step professional tasks the model must carry out in a command-line interface. Reported as pass@1: one attempt per task.
- FrontierCode 1.1Agentic coding
- Cognition's agentic coding test: the model resolves real issues in open-source repositories, and each patch is graded by hidden tests and a maintainer-written rubric. Main is the 100 hardest tasks.
- ChartographyVisual reasoning
- Visual chart recognition: reading values and trends from charts and figures. Haiku 5.5's score is without tools.
Haiku 5.5 benchmarks: questions
How good is Claude Haiku 5.5 on benchmarks?
It is the strongest Haiku yet: 72.4% on OSWorld 2.1, 39.2% on Terminal-Bench 4.0, 57.4% on Humanity's Last Exam with tools and 1620 Elo on GDPval-AA. In Anthropic's table it beats GPT-6 Luna on every benchmark where Luna has a score, and trails Sonnet 5.5 on all of them.
What is Haiku 5.5's SWE-bench score?
Anthropic's launch post doesn't list SWE-bench. Launch coverage of the Haiku 5.5 system card reports 64.8% on SWE-bench Pro and 83.7% on SWE-bench Multilingual.
Is Haiku 5.5 better than Haiku 4.5?
By a wide margin on every published benchmark, for example 72.4% vs 15.7% on OSWorld 2.1 and 39.2% vs 0.0% on Terminal-Bench 4.0, while costing up to 90% less.
Is Haiku 5.5 better than GPT-6 Luna?
On Anthropic's table, yes: it leads on all six shared benchmarks. Independent testing by Artificial Analysis puts Haiku 5.5 ahead at max effort (43 vs 38 on its Intelligence Index) and level with Luna at high effort.
How does Haiku 5.5 compare with Opus 5.5?
Opus 5.5 leads on every benchmark both launch posts report, for example 66.4% vs 39.2% on Terminal-Bench 4.0, but costs 40× more per token. Our Haiku 5.5 vs Opus 5.5 page has the full comparison.
Keep reading
- Haiku 5.5 vs Opus 5.5Benchmarks, price, speed and when to use each.
- Haiku 5.5 pricingThe full rate card, the 100K-token rule and a cost calculator.
- Haiku 5.5 vs Haiku 4.5What changed in the new Haiku, including API breaking changes.
- Chat with Haiku 5.5 freeNo sign-up needed. Web search and link reading built in.
Try Claude Haiku 5.5 free, right now
No credit card, nothing to install. Ask your first question in seconds.
Start chatting freeSources
- Anthropic — Introducing Claude Haiku 5.5
- Anthropic — Claude Haiku 5.5 system card
- Anthropic — Claude Haiku (availability, customer results)
- Anthropic — Introducing Claude Opus 5.5
- Artificial Analysis — Claude Haiku 5.5
- The Decoder — Claude Haiku 5.5 launch coverage
- Handy AI — Model Drop: Claude Haiku 5.5 (system card figures)
Haiku55 is an independent service and is not affiliated with Anthropic. Prices and scores are as published by Anthropic and the testers named above, checked on October 8, 2026; see the sources for later changes.