Anthropic Ships Claude Haiku 5.5; Independent Tests Find Smaller Agentic-Coding Gain Than Claimed

Anthropic's cheapest model gained a 1M-token context window and sharply higher scores on October 7, 2026, but Artificial Analysis measured a smaller agentic-coding gain than Anthropic reported.

EduFabTech Β· 10 October 2026 Β· 3 min read Β· 1 views
Haiku 5.5's 1M-token context window and Anthropic's reported OSWorld and Terminal-Bench gains, alongside the lower Terminal-Bench score Artificial Analysis independently measured.
EduFabTech · Own work

Anthropic released Claude Haiku 5.5 on October 7, 2026, replacing Haiku 4.5 as the cheapest model in the Claude lineup and giving it a 1-million-token context window, up from 200,000 tokens, according to Anthropic's release page. The company's own benchmark figures show large jumps on tasks involving computer use and command-line coding, but independent testing by Artificial Analysis, published the same week, measured a smaller gain on one of those same benchmarks β€” a gap Anthropic has not addressed publicly.

Haiku has always been positioned as Anthropic's budget tier: the model meant for high-volume, repetitive work rather than frontier reasoning. Anthropic describes Haiku 5.5 as built for tasks like "summaries, compactions, database queries, and classification requests," and as a subagent that larger models like Sonnet 5.5 can delegate routine steps to, according to the same release page. The model also adds adaptive thinking and an adjustable "effort" setting, letting a developer trade speed for quality without switching models, detailed in Anthropic's Claude Platform documentation.

A bar chart of Terminal-Bench 4.0 scores across Haiku 4.5, GPT-6 Luna, Gemini 3.8 Flash, GLM-5.3 Flash, and Haiku 5.5, with a dashed line marking Anthropic's higher self-reported Haiku 5.5 figure above the bar Artificial Analysis measured.
A bar chart of Terminal-Bench 4.0 scores across Haiku 4.5, GPT-6 Luna, Gemini 3.8 Flash, GLM-5.3 Flash, and Haiku 5.5, with a dashed line marking Anthropic's higher self-reported Haiku 5.5 figure above the bar Artificial Analysis measured.EduFabTech · Own work

What Anthropic's own numbers show

On the OSWorld 2.1 offline subset, a benchmark that scores an AI agent's ability to operate a computer desktop, Anthropic reports Haiku 5.5 reaching 72.4%, up from 15.7% for Haiku 4.5 and ahead of OpenAI's GPT-6 Luna at 48.9%, according to Anthropic's release page. On Terminal-Bench 4.0, which tests an agent's ability to complete tasks in a command-line terminal, Anthropic reports Haiku 5.5 at 39.2%, up from 0.0% for Haiku 4.5. On the GDPval-AA v2.1 knowledge benchmark, which reports Elo-style ratings rather than percentages, Haiku 5.5 scores 1,620 against Haiku 4.5's 735.

Pricing follows a tiered structure: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising to $0.50 and $2.50 per million above that threshold, per the Claude Platform documentation. Anthropic says this amounts to roughly 75% lower cost than Haiku 4.5 for typical workloads.

Where the independent numbers diverge

Artificial Analysis, a third-party benchmarking organization that runs its own test harnesses rather than relying on vendor-reported scores, published its evaluation of Haiku 5.5 in an article the same week, placing the model at 43 points on its Intelligence Index at maximum effort β€” a 26-point rise over Haiku 4.5 β€” ranking it ahead of GLM-5.3 Flash (42) and Gemini 3.8 Flash (41), behind Kimi K3 (44), and well below Claude Sonnet 5.5 at maximum effort (56), according to Artificial Analysis's writeup.

On Terminal-Bench 4.0 specifically, Artificial Analysis measured Haiku 5.5 at 33%, below Anthropic's own reported 39.2%, though still a sharp rise from Haiku 4.5's 0%. Artificial Analysis's figure put Haiku 5.5 level with GLM-5.3 Flash and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%) on the same test, per the same writeup. Neither organization has published an explanation for the gap between the two Terminal-Bench figures.

Side-by-side panels contrasting Anthropic's self-reported benchmark figures with Artificial Analysis's independently measured scores and token-efficiency numbers for Haiku 5.5.
Side-by-side panels contrasting Anthropic's self-reported benchmark figures with Artificial Analysis's independently measured scores and token-efficiency numbers for Haiku 5.5.EduFabTech · Own work

A cost tradeoff behind the headline score

Artificial Analysis also flagged an efficiency difference the headline comparisons omit: at maximum effort, Haiku 5.5 consumes approximately 162,000 output tokens per Intelligence Index task, about three times the roughly 50,000 tokens GPT-6 Luna uses for a comparable score, according to the Artificial Analysis article. Because Haiku 5.5 is also billed by output token, a higher benchmark score at maximum effort does not automatically translate into a cheaper result in practice β€” the actual cost depends on how many tokens a given task burns through, not just the per-million-token rate.

What it means for people building with the model

For researchers and engineers choosing among small, cheap models for high-volume pipelines, the practical takeaway is that vendor-reported scores and independently measured scores can diverge even on an open, nameable benchmark like Terminal-Bench 4.0 β€” and that the gap can run in either direction depending on effort level and harness. Anyone selecting a model for an agentic coding or computer-use workload should check the effort setting and token-consumption figures reported alongside any benchmark score, not just the headline percentage, since Artificial Analysis's own numbers show the same model scoring differently depending on how much it is allowed to "think" before answering.

Anthropic's and Artificial Analysis's benchmark figures were published within the same week, and neither organization has yet issued a reconciliation of the Terminal-Bench difference.

A quick question for readers

Comments 0

Comment

No comments yet. Members start the conversation.

Source: Anthropic

Sources (3)
  1. Anthropic. Introducing Claude Haiku 5.5. Anthropic, 2026. anthropic.com β†— Β· checked 10 Oct 2026
  2. Artificial Analysis. Anthropic has released Claude Haiku 5.5. Artificial Analysis, 2026. artificialanalysis.ai β†— Β· checked 10 Oct 2026
  3. Anthropic. What's new in Claude Haiku 5.5. Claude Platform Docs, 2026. platform.claude.com β†— Β· checked 10 Oct 2026