Artificial Analysis has released its evaluation of Claude Haiku 5.5, giving the model a composite intelligence score of 43. That marks a 26-point jump from the previous generation’s 17 and puts it ahead of GPT-6 Luna at 38 and GLM-5.3 Flash at 42, while sitting close to Kimi K3 at 44. The benchmark combines 10 tests covering areas including coding, professional work, and scientific reasoning.
The higher score came with a steep token bill. At the highest reasoning setting, Haiku 5.5 produced about 162,000 tokens per task on average, or 3.2 times Luna’s output. Although the two models share the same base API pricing, Artificial Analysis estimated Haiku 5.5’s cost per task at about $0.21, versus $0.07 for Luna. Moving Haiku 5.5 from the second-highest setting to the top setting added only 2 points in score while increasing output by roughly 80%.
At the high setting, the gap narrowed: Haiku 5.5 scored 38 with about 55,000 tokens per task, close to Luna’s 38 and 50,000 tokens at its highest setting. In AutomationBench-AA, however, Haiku 5.5 scored 35%, trailing Luna’s 53%. Artificial Analysis said the model showed excessive refusal behavior during testing. Anthropic is working on a fix, and the firm plans to run the test again. The published task-cost estimate also does not yet include long-context surcharges.
Third-party evaluator Artificial Analysis has published test results for Claude Haiku 5.5, giving the model a composite intelligence score of 43. That was up 26 points from the previous generation’s 17.
On the same index, Haiku 5.5 ranked above GPT-6 Luna at 38 and GLM-5.3 Flash at 42, and came close to Kimi K3 at 44. Artificial Analysis said the index combines 10 tests, including coding, professional work, and scientific reasoning.
Higher scores came with heavier token usage
At the highest reasoning intensity, Haiku 5.5 generated about 162,000 tokens per evaluation task on average, 3.2 times the level of GPT-6 Luna. The two models have the same base API pricing, but Haiku 5.5’s estimated cost per task was about $0.21, compared with $0.07 for Luna.
Artificial Analysis also found that moving Haiku 5.5 from the second-highest setting to the top setting lifted the composite score by only 2 points, while output volume rose by about 80%.
The gap narrowed at lower reasoning intensity
At the high setting, Haiku 5.5 scored 38 and produced about 55,000 tokens per task on average. That was close to GPT-6 Luna’s 38 and 50,000 tokens at its highest setting. The results suggest developers can reduce unnecessary token consumption by adjusting reasoning intensity.
Automation task results lagged Luna
Haiku 5.5 did not lead across every benchmark. In the AutomationBench-AA automation task evaluation, it scored 35%, behind Luna’s 53%.
Artificial Analysis said the model showed excessive refusal to carry out tasks during the test period. Anthropic is fixing the issue, and Artificial Analysis plans to test the model again. The currently published task-cost figure for Haiku 5.5 also does not include long-context surcharges, meaning actual costs could be higher.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.