Benchmarks · updated October 8, 2026

Haiku 5.5 benchmark results

Anthropic launched Claude Haiku 5.5 on October 7, 2026 with eight headline benchmark scores. Here is every one of them next to Haiku 4.5, Sonnet 5.5 and OpenAI's GPT-6 Luna: what each test measures, how big the jump is, and what independent testers found.

OSWorld 2.1, computer use. Haiku 4.5: 15.7%
72.4%
Terminal-Bench 4.0, agentic coding. Haiku 4.5: 0.0%
39.2%
Humanity's Last Exam with tools. Haiku 4.5: 18.7%
57.4%
GDPval-AA Elo, knowledge work. Haiku 4.5: 735
1620

All results at a glance

Anthropic's published table, unchanged. Higher is better on every row. GDPval-AA and AA-Briefcase are Elo ratings; the rest are percentages.

Claude Haiku 5.5 benchmark scores compared with Haiku 4.5, GPT-6 Luna and Sonnet 5.5
BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1Knowledge work · Elo rating162073514371840
AA-Briefcase v1.1Knowledge work · Elo rating157861413361824
OSWorld 2.1Computer use · offline subset72.4%15.7%48.9%83.9%
Humanity's Last ExamReasoning · no tools45.9%10.2%Not reported56.9%
Humanity's Last ExamReasoning · with tools57.4%18.7%Not reported64.5%
Terminal-Bench 4.0Agentic coding · pass@139.2%0.0%16.4%70.6%
FrontierCode 1.1Agentic coding · Main46.4%Not reported42.4%52.1%xhigh effort
ChartographyVisual reasoning · no tools46.4%6.4%29.1%61.6%

Source: Anthropic, “Introducing Claude Haiku 5.5”, October 7, 2026. A dash means not reported. Sonnet 5.5's FrontierCode score was run at xhigh effort. Evaluation details are in the Haiku 5.5 system card.

Haiku 5.5 vs Haiku 4.5: a generational jump

Haiku 4.5 was a fast chat model; Haiku 5.5 is a working agent. The biggest gains are on tests where the model has to act, not just answer.

GDPval-AA v2.1

Knowledge work · Elo rating

+885Elo

Haiku 5.5
1620
Haiku 4.5
735

AA-Briefcase v1.1

Knowledge work · Elo rating

+964Elo

Haiku 5.5
1578
Haiku 4.5
614

OSWorld 2.1

Computer use · offline subset

4.6×higher

Haiku 5.5
72.4%
Haiku 4.5
15.7%

Humanity's Last Exam

Reasoning · no tools

4.5×higher

Haiku 5.5
45.9%
Haiku 4.5
10.2%

Humanity's Last Exam

Reasoning · with tools

3.1×higher

Haiku 5.5
57.4%
Haiku 4.5
18.7%

Terminal-Bench 4.0

Agentic coding · pass@1

+39.2points

Haiku 5.5
39.2%
Haiku 4.5
0.0%

Chartography

Visual reasoning · no tools

7.3×higher

Haiku 5.5
46.4%
Haiku 4.5
6.4%

On Terminal-Bench 4.0, which has the model complete multi-step professional tasks in a command line, Haiku 4.5 scored 0.0% in Anthropic's run and Haiku 5.5 scores 39.2%. On OSWorld 2.1, where it operates a real computer through long tasks, it climbs from 15.7% to 72.4%. And it gets there at up to 90% lower cost.

Haiku 5.5 vs GPT-6 Luna

GPT-6 Luna is OpenAI's small model and the rival Anthropic chose to compare against. Haiku 5.5 leads on all six benchmarks where both have a score.

GDPval-AA v2.1

Knowledge work · Elo rating

+183Elo ahead of Luna

Haiku 5.5
1620
GPT-6 Luna
1437

AA-Briefcase v1.1

Knowledge work · Elo rating

+242Elo ahead of Luna

Haiku 5.5
1578
GPT-6 Luna
1336

OSWorld 2.1

Computer use · offline subset

+23.5points ahead of Luna

Haiku 5.5
72.4%
GPT-6 Luna
48.9%

Terminal-Bench 4.0

Agentic coding · pass@1

+22.8points ahead of Luna

Haiku 5.5
39.2%
GPT-6 Luna
16.4%

FrontierCode 1.1

Agentic coding · Main

+4.0points ahead of Luna

Haiku 5.5
46.4%
GPT-6 Luna
42.4%

Chartography

Visual reasoning · no tools

+17.3points ahead of Luna

Haiku 5.5
46.4%
GPT-6 Luna
29.1%

Independent testing is closer. Artificial Analysis gives Haiku 5.5 a 43 on its Intelligence Index at max effort, ahead of GPT-6 Luna's 38. At high effort the two tie at 38, and at max effort Haiku 5.5 used about three times as many output tokens as Luna. Artificial Analysis also measured a lower hallucination rate for Haiku 5.5 (40% vs 77%) and higher knowledge accuracy for Luna (44% vs 36%).

How close is Haiku 5.5 to Sonnet 5.5?

Sonnet 5.5 stays ahead on every benchmark. The real question is how much of Sonnet's score you get for 1/20 of its per-token price.

OSWorld 2.1

Computer use · offline subset

86%of Sonnet 5.5's score

Haiku 5.5
72.4%
Sonnet 5.5
83.9%

Humanity's Last Exam

Reasoning · no tools

81%of Sonnet 5.5's score

Haiku 5.5
45.9%
Sonnet 5.5
56.9%

Humanity's Last Exam

Reasoning · with tools

89%of Sonnet 5.5's score

Haiku 5.5
57.4%
Sonnet 5.5
64.5%

Terminal-Bench 4.0

Agentic coding · pass@1

56%of Sonnet 5.5's score

Haiku 5.5
39.2%
Sonnet 5.5
70.6%

FrontierCode 1.1

Agentic coding · Main

89%of Sonnet 5.5's score

Haiku 5.5
46.4%
Sonnet 5.5
52.1%

Chartography

Visual reasoning · no tools

75%of Sonnet 5.5's score

Haiku 5.5
46.4%
Sonnet 5.5
61.6%

On reasoning, computer-use and graded coding tests, Haiku 5.5 reaches 81–89% of Sonnet 5.5's score. On long agentic coding the gap is real: 39.2% vs 70.6% on Terminal-Bench 4.0. That matches Anthropic's own advice: Sonnet 5.5 or Opus 5.5 for complex agentic coding, and Haiku 5.5 for high-volume work and as their sub-agent.

What about SWE-bench?

Anthropic's launch post does not include SWE-bench; its coding benchmarks are Terminal-Bench 4.0 and FrontierCode. Launch coverage citing the Haiku 5.5 system card reports 64.8% on SWE-bench Pro and 83.7% on SWE-bench Multilingual.

You may also see Haiku 4.5 quoted at 73.3%. That score is on SWE-bench Verified, a different test set, so it can't be compared with the numbers above.

Effort changes the score

Haiku 5.5 is the first Haiku with effort levels: low, medium (the default), high, xhigh and max. Anthropic charts each level's accuracy against its cost on OSWorld, GDPval-AA and Humanity's Last Exam. Scores climb with effort, and so does token use.

Artificial Analysis saw the same curve: Haiku 5.5 scores 38 on its Intelligence Index at high effort and 43 at max, and the step from xhigh to max adds two points for about 1.8× the tokens. For everyday chat, low or medium effort is the sweet spot; save the top levels for hard problems. On Haiku55, the Fast, Smart and Deep modes map to low, medium and high effort.

What early testers measured

Results that companies reported from their own evaluations at launch:

  • HubSpot

    92.8% on its CRM simulation suite, averaged over three runs: the best score HubSpot has seen on it.

  • AlphaSense

    0.84 versus 0.76 for Haiku 4.5 across 400 queries on its Ask in Document feature.

  • Box

    11 points higher than Haiku 4.5 at about half the latency.

  • Asana

    Over 30% lower task-completion latency and up to 2.5× faster inference per agent turn.

What each benchmark measures

GDPval-AA v2.1Knowledge work
Real professional work drawn from 44 occupations. Models work as agents, and blind head-to-head grading ranks them on an Elo scale. Run by Artificial Analysis.
AA-Briefcase v1.1Knowledge work
Artificial Analysis's private test of long-horizon business work, where agents produce deliverables such as spreadsheets, presentations and memos. Also scored as Elo.
OSWorld 2.1Computer use
The model operates a real computer to finish long, multi-step tasks across desktop apps. Anthropic reports Haiku 5.5 on the offline subset.
Humanity's Last ExamReasoning
Humanity's Last Exam: expert-level questions across dozens of academic fields, written to be hard for frontier models. Reported both with and without tools.
Terminal-Bench 4.0Agentic coding
Multi-step professional tasks the model must carry out in a command-line interface. Reported as pass@1: one attempt per task.
FrontierCode 1.1Agentic coding
Cognition's agentic coding test: the model resolves real issues in open-source repositories, and each patch is graded by hidden tests and a maintainer-written rubric. Main is the 100 hardest tasks.
ChartographyVisual reasoning
Visual chart recognition: reading values and trends from charts and figures. Haiku 5.5's score is without tools.

Haiku 5.5 benchmarks: questions

How good is Claude Haiku 5.5 on benchmarks?

It is the strongest Haiku yet: 72.4% on OSWorld 2.1, 39.2% on Terminal-Bench 4.0, 57.4% on Humanity's Last Exam with tools and 1620 Elo on GDPval-AA. In Anthropic's table it beats GPT-6 Luna on every benchmark where Luna has a score, and trails Sonnet 5.5 on all of them.

What is Haiku 5.5's SWE-bench score?

Anthropic's launch post doesn't list SWE-bench. Launch coverage of the Haiku 5.5 system card reports 64.8% on SWE-bench Pro and 83.7% on SWE-bench Multilingual.

Is Haiku 5.5 better than Haiku 4.5?

By a wide margin on every published benchmark, for example 72.4% vs 15.7% on OSWorld 2.1 and 39.2% vs 0.0% on Terminal-Bench 4.0, while costing up to 90% less.

Is Haiku 5.5 better than GPT-6 Luna?

On Anthropic's table, yes: it leads on all six shared benchmarks. Independent testing by Artificial Analysis puts Haiku 5.5 ahead at max effort (43 vs 38 on its Intelligence Index) and level with Luna at high effort.

How does Haiku 5.5 compare with Opus 5.5?

Opus 5.5 leads on every benchmark both launch posts report, for example 66.4% vs 39.2% on Terminal-Bench 4.0, but costs 40× more per token. Our Haiku 5.5 vs Opus 5.5 page has the full comparison.

Try Claude Haiku 5.5 free, right now

No credit card, nothing to install. Ask your first question in seconds.

Start chatting free