Anthropic Ships Claude Haiku 5.5; Independent Tests Find Smaller Agentic-Coding Gain Than Claimed
Anthropic's cheapest model gained a 1M-token context window and sharply higher scores on October 7, 2026, but Artificial Analysis measured a smaller agentic-coding gain than Anthropic reported.
Anthropic released Claude Haiku 5.5 on October 7, 2026, replacing Haiku 4.5 as the cheapest model in the Claude lineup and giving it a 1-million-token context window, up from 200,000 tokens, according to Anthropic's release page. The company's own benchmark figures show large jumps on tasks involving computer use and command-line coding, but independent testing by Artificial Analysis, published the same week, measured a smaller gain on one of those same benchmarks β a gap Anthropic has not addressed publicly.
Haiku has always been positioned as Anthropic's budget tier: the model meant for high-volume, repetitive work rather than frontier reasoning. Anthropic describes Haiku 5.5 as built for tasks like "summaries, compactions, database queries, and classification requests," and as a subagent that larger models like Sonnet 5.5 can delegate routine steps to, according to the same release page. The model also adds adaptive thinking and an adjustable "effort" setting, letting a developer trade speed for quality without switching models, detailed in Anthropic's Claude Platform documentation.

What Anthropic's own numbers show
On the OSWorld 2.1 offline subset, a benchmark that scores an AI agent's ability to operate a computer desktop, Anthropic reports Haiku 5.5 reaching 72.4%, up from 15.7% for Haiku 4.5 and ahead of OpenAI's GPT-6 Luna at 48.9%, according to Anthropic's release page. On Terminal-Bench 4.0, which tests an agent's ability to complete tasks in a command-line terminal, Anthropic reports Haiku 5.5 at 39.2%, up from 0.0% for Haiku 4.5. On the GDPval-AA v2.1 knowledge benchmark, which reports Elo-style ratings rather than percentages, Haiku 5.5 scores 1,620 against Haiku 4.5's 735.
Pricing follows a tiered structure: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising to $0.50 and $2.50 per million above that threshold, per the Claude Platform documentation. Anthropic says this amounts to roughly 75% lower cost than Haiku 4.5 for typical workloads.
Where the independent numbers diverge
Artificial Analysis, a third-party benchmarking organization that runs its own test harnesses rather than relying on vendor-reported scores, published its evaluation of Haiku 5.5 in an article the same week, placing the model at 43 points on its Intelligence Index at maximum effort β a 26-point rise over Haiku 4.5 β ranking it ahead of GLM-5.3 Flash (42) and Gemini 3.8 Flash (41), behind Kimi K3 (44), and well below Claude Sonnet 5.5 at maximum effort (56), according to Artificial Analysis's writeup.
On Terminal-Bench 4.0 specifically, Artificial Analysis measured Haiku 5.5 at 33%, below Anthropic's own reported 39.2%, though still a sharp rise from Haiku 4.5's 0%. Artificial Analysis's figure put Haiku 5.5 level with GLM-5.3 Flash and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%) on the same test, per the same writeup. Neither organization has published an explanation for the gap between the two Terminal-Bench figures.

A cost tradeoff behind the headline score
Artificial Analysis also flagged an efficiency difference the headline comparisons omit: at maximum effort, Haiku 5.5 consumes approximately 162,000 output tokens per Intelligence Index task, about three times the roughly 50,000 tokens GPT-6 Luna uses for a comparable score, according to the Artificial Analysis article. Because Haiku 5.5 is also billed by output token, a higher benchmark score at maximum effort does not automatically translate into a cheaper result in practice β the actual cost depends on how many tokens a given task burns through, not just the per-million-token rate.
What it means for people building with the model
For researchers and engineers choosing among small, cheap models for high-volume pipelines, the practical takeaway is that vendor-reported scores and independently measured scores can diverge even on an open, nameable benchmark like Terminal-Bench 4.0 β and that the gap can run in either direction depending on effort level and harness. Anyone selecting a model for an agentic coding or computer-use workload should check the effort setting and token-consumption figures reported alongside any benchmark score, not just the headline percentage, since Artificial Analysis's own numbers show the same model scoring differently depending on how much it is allowed to "think" before answering.
Anthropic's and Artificial Analysis's benchmark figures were published within the same week, and neither organization has yet issued a reconciliation of the Terminal-Bench difference.
A quick question for readers
Which skill will matter most for getting a job in 2030?
Everyone has a prediction. Add yours and see what students are betting on.
Pick an answer, create a free account in a minute, and your vote counts. Already a member? Sign in
Comments 0
CommentNo comments yet. Members start the conversation.
Source: Anthropic
Sources (3)
- Anthropic. Introducing Claude Haiku 5.5. Anthropic, 2026. anthropic.com β Β· checked 10 Oct 2026
- Artificial Analysis. Anthropic has released Claude Haiku 5.5. Artificial Analysis, 2026. artificialanalysis.ai β Β· checked 10 Oct 2026
- Anthropic. What's new in Claude Haiku 5.5. Claude Platform Docs, 2026. platform.claude.com β Β· checked 10 Oct 2026