Anthropic launches Claude Haiku 5.5 with API pricing identical to OpenAI’s GPT-6 Luna

Anthropic launches Claude Haiku 5.5 with API pricing identical to OpenAI’s GPT-6 Luna

N
News Editor
2026-10-08 03:21:31
Anthropic on Oct. 7 launched Claude Haiku 5.5 across Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, positioning it as a low-cost, high-speed small model for subagent workloads. Its four API price points match OpenAI’s GPT-6 Luna exactly: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 for cache reads, and $0.125 for cache writes. Anthropic said short requests can be 90% cheaper than Haiku 4.5, though that headline figure comes with limits. Requests above 100,000 tokens are priced at $0.50 for input and $2.50 for output, which cuts the discount to 50%. The company also said its new tokenizer counts roughly 30% more tokens than Haiku 4.5 for the same text, putting the average price reduction cited in the announcement at about 75%. Anthropic’s benchmark table was published using the Max reasoning setting rather than the default Medium level, and scores drop materially at the default setting. The company said Haiku 5.5 is meant to support larger models as a subagent for summarization, context compression, database lookup, and classification, while more complex agentic coding tasks should still be handled by Sonnet 5.5 and Opus 5.5.

Claude Haiku 5.5 goes live across four platforms

Anthropic launched Claude Haiku 5.5 on Oct. 7 U.S. time, making the model available on Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure.

Its API pricing matches OpenAI’s GPT-6 Luna across all four listed categories: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 for cache reads, and $0.125 for cache writes.

Google’s Gemini 3.5 Flash-Lite remains priced higher on the standard list, at $0.30 for input and $2.50 for output. That leaves its output price at five times the level listed by Anthropic and OpenAI. Still, identical list prices do not automatically produce identical bills.

The 90% discount applies mainly to shorter requests

The headline number in Anthropic’s release is a 90% price cut, but the company tied that figure to request length.

According to Anthropic’s pricing page, prompts are split into two tiers. For requests at or below 100,000 tokens, both input and output pricing are one-tenth of Haiku 4.5. Above 100,000 tokens, pricing rises to $0.50 for input and $2.50 for output, which means the discount narrows to 50%. The gap between the two tiers is fivefold.

Anthropic said about 90% of Haiku 4.5 requests fall within the 100,000-token range, so most requests qualify for the steeper discount.

There is another adjustment. Haiku 5.5 uses a new tokenizer, meaning the same text is split into tokens under a different set of rules. In Anthropic’s update notes, the company said the same passage will be counted as roughly 30% more tokens than under Haiku 4.5. With lower unit pricing, a higher token count, and around one-tenth of long requests receiving only a 50% cut, Anthropic said the average reduction in the body of its announcement is about 75%.

Benchmark scores were published at Max effort, not the default setting

Anthropic also highlighted a separate cost variable: effort, or reasoning intensity. The company lists five levels — Low, Medium, High, Xhigh, and Max — with Medium set as the API default.

The benchmark table in Anthropic’s announcement shows Haiku 5.5 results at Max. Those figures were published by Anthropic itself and were not described as third-party validation.

At the default Medium setting, several benchmark results fall sharply. On OSWorld, a computer-use benchmark, the score drops from 72.4% to 53.3%. On GDPval-AA, a knowledge-work benchmark, it falls from 1620 to 1277. On Terminal-Bench, which measures agentic coding, the score declines from 39.2% to 20.3%.

Anthropic still points complex coding work to larger models

Coding is a separate case in Anthropic’s own framing. On Terminal-Bench, Haiku 5.5 Max scored 39.2%, but its cost was about 80% higher than Sonnet 5.5 High at 43%.

Anthropic said directly in the announcement that complex agentic coding tasks should still be handled by Sonnet 5.5 and Opus 5.5.

Haiku 5.5 is positioned as a subagent model

Anthropic’s positioning for Haiku 5.5 is straightforward. The model is not presented as a replacement for larger models. It is meant to act as a subagent that handles smaller, clearly scoped tasks delegated by a larger model.

Anthropic listed summarization, context compression, database lookup, and classification among the intended use cases, describing them as high-volume jobs where cost matters. On speed, the company called Haiku 5.5 its fastest model at standard speed, while a footnote said it is still slower than Opus in Fast Mode. The announcement did not provide a tokens-per-second figure.

Customer examples in the announcement came from Rogo and AlphaSense

Anthropic cited financial AI company Rogo as one example. In that workflow, a larger model prepares a presentation, while a Haiku 5.5 subagent goes into a 10-K filing and retrieves the line item for segment revenue needed in the slides. Alex Wang of Rogo said, “It’s accurate enough that we’re comfortable trusting it, and fast and cheap enough to run at scale.”

Anthropic also cited AlphaSense. The company said its Ask in Document feature handles about 8 million calls per week. In a test covering 400 queries, Haiku 5.5 scored 0.84, compared with 0.76 for Haiku 4.5.

Both customer references appeared in Anthropic’s announcement.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.