Anthropic has released Claude Haiku 5.5, pricing it at one-tenth the cost of the previous Haiku 4.5 and one-twentieth the price of Sonnet 5.5. The company’s pricing also matches OpenAI’s GPT-6 Luna for requests under 100,000 tokens.
In Anthropic’s published benchmark table, Haiku 5.5 outperformed Luna in all six tests where Luna submitted results. In several categories, it also closed in on Sonnet 5.5.
Pricing and cost profile
For prompts up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. For requests above 100,000 tokens, pricing rises to $0.50 for input and $2.50 for output per million tokens.
Cache reads are priced at $0.01 per million tokens, and batch processing cuts that in half again.
Anthropic said its internal estimate shows the same workload is about 75% cheaper on average than Haiku 4.5, which means roughly four times as many calls under the same budget.
The article noted that the overall savings do not reach 90% because requests above 100,000 tokens are only half as cheap. It also said roughly 90% of Haiku requests in the previous generation stayed within the 100,000-token range, while the new tokenizer adds about 30% more tokens for the same passage of text.
Against GPT-6 Luna, input, output, and cache pricing are identical below 100,000 tokens. In long-context use cases, Luna is cheaper. Its price does not increase until around 270,000 tokens, and after that point it rises only to $0.20 for input and $0.75 for output.
Anthropic also cut Sonnet 5.5 cache read pricing in half starting the same day, from $0.20 to $0.10 per million tokens. The company said that change lowers costs for most agent tasks by about 20%.
On subscriptions, Max and Team plans now come with monthly API credits starting this week. Max 5x includes $100, Max 20x includes $200, and Team includes up to $500 shared across the team. The credits can be used for any model on Claude Platform.
Benchmarks improved sharply, but complex coding still trails Sonnet
Haiku 5.5 posted large gains over the previous generation across a wide set of benchmarks.
In OSWorld, which measures computer-use performance, the model jumped from 15.7% to 72.4%, leaving it 11.5 percentage points behind Sonnet 5.5. Luna scored 48.9% in the same test, so Haiku 5.5 led by 23.5 percentage points. The article also said that on a cost-adjusted basis, Haiku 5.5 at the Medium setting outperformed Luna at its highest setting while spending less.
On Terminal-Bench 4.0, Haiku 4.5 previously scored zero. Haiku 5.5 reached 39.2%. That was well above Luna’s 16.4%, but still far below Sonnet 5.5 at 70.6%.
On GDPval-AA, a benchmark for knowledge work, Haiku 5.5 scored 1,620, more than double the previous generation. That result beat Luna’s 1,437 but remained behind Sonnet 5.5 at 1,840. At the default Medium setting, the score was 1,277.
On HLE without tools, described in the article as “Humanity’s Last Exam,” Haiku 5.5 improved from 10.2% to 45.9%. With search and coding tools enabled, Haiku 5.5 reached 57.4%, versus 64.5% for Sonnet 5.5.
On AA-Briefcase, which the article described as a benchmark for long-running projects, Haiku 5.5 rose from 614 to 1,578. Luna scored 1,336.
On Chartography, a benchmark for reading specialized charts, Haiku 5.5 improved from 6.4% to 46.4%. Luna posted 29.1%, while Sonnet 5.5 reached 61.6%.
On FrontierCode, Haiku 5.5 scored 46.4%, ahead of Luna’s 42.4% and about 6 percentage points behind Sonnet 5.5’s best result of 52.1%.
Anthropic also said complex agent coding is still better handled by Sonnet 5.5 and Opus 5.5.
Positioned for fast, high-volume workloads
Anthropic described Haiku 5.5 as a model for tasks that need scale, speed, and lower cost, including summaries, context compression, database queries, and classification.
The article said it is currently the fastest Claude model, slower only than Opus running in fast mode, which it described as expensive. That makes it suitable for live customer support and browser-based actions.
Anthropic also shared early feedback from AI application companies.
According to the article, Asana found that task completion latency was cut by more than 30% compared with the model it currently uses. Box said latency was about half that of Haiku 4.5.
Another use case is as a sub-agent under Opus or Sonnet, where the larger model breaks down tasks and makes decisions while Haiku handles execution in parallel.
Cognition’s Devin Fusion used Opus 5.5 as the lead model and Haiku 5.5 as support. The article said FrontierCode stayed at a top-tier 66.2 points while both cost and latency fell.
Financial AI company Rogo used the model in a slide-generation workflow, with Haiku 5.5 extracting segment revenue data from 10-K filings.
The Claude developer account also showed a demo in which the system had to design an egg-protection rack from household items that could survive a 32-meter drop. Opus 5.5 alone took 3 minutes and 37 seconds, tested 25 designs, and cost $0.47. With 10 Haiku 5.5 sub-agents added, the job finished in 58 seconds, tested 86 designs, and cost $0.14.
Availability and migration changes
Haiku 5.5 is now available in the Claude app, Claude Code, AWS, Google Cloud, Microsoft Azure, Cursor, and OpenRouter. The model ID is claude-haiku-5-5. It supports a 1 million-token context window and up to 128,000 output tokens in a single response.
It is the first Haiku model with adjustable thinking levels, offering five settings from Low to Max. The API default is Medium.
The Python and TypeScript SDKs now include computer-use and browser-use capabilities in beta.
For developers migrating from Haiku 4.5, Anthropic said the same text now produces about 30% more tokens, so max_tokens settings and cost estimates need to be recalculated.
On the API side, budget_tokens has been replaced by adaptive thinking with Effort levels. temperature, top_p, and top_k must be removed. Prefill responses are no longer supported, and formatting requirements now use structured output.
Computer-use features must move to the new toolset, computer_toolset_20260801. Priority Tier is not supported for now.
The article said developers can also run a command in Claude Code to let the system handle the migration. Non-developers, meanwhile, can start using Haiku 5.5 directly in the client.

