Alibaba unveils Qwen3.8-Max with self-reported benchmark lead and aggressive token pricing

Alibaba unveils Qwen3.8-Max with self-reported benchmark lead and aggressive token pricing

N
News Editor
2026-08-04 01:51:01
Alibaba’s Qwen team has introduced Qwen3.8-Max, a new flagship model that the company says scored 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2, Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2. The release also included a broader slate of benchmark claims, such as 93.0 on PaperBench and 86.6 on TerminalBench 2.1, alongside positioning the model for long-running autonomous work rather than standard chatbot use. The model uses a mixture-of-experts architecture with 2.4 trillion total parameters and about 95 billion active during inference, built on the Qwen3.5 architecture with a 1 million-token context window. Alibaba also said Qwen3.8-Max is suited for extended coding tasks, desktop software operation, experiment reproduction, and industrial workflows that feed visual input back into a planning loop. Pricing appears to be a central part of the launch. According to QwenCloud pricing cited in the report, Qwen3.8-Max costs $2 per million input tokens and $6 per million output tokens overseas, bringing the combined total to $8 per million tokens. That is below one-third of Claude Opus 5’s combined $30 and below one-quarter of GPT-5.6 Sol standard mode at $35. Still, the benchmarks and capability demonstrations were all disclosed by Alibaba and have not been independently verified, while the company has yet to publish the licensing terms for the promised open-weight release next week on Hugging Face and ModelScope.

Alibaba’s Qwen team has launched Qwen3.8-Max, a new flagship model that the company says scored 86.1 on OSWorld-Verified, a benchmark for computer-use agent capability that tests whether a model can operate desktop software in a human-like way. In Alibaba’s published comparison, that put Qwen3.8-Max ahead of GPT-5.6 Sol Max at 83.2, Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2.

The report, originally published by BlockTempo, noted that the scorecard currently has only one evaluator: Alibaba itself. The benchmark figures released for Qwen3.8-Max have not been independently verified by a third party.

Beyond OSWorld-Verified, Alibaba listed several additional results for the model, including 93.0 on PaperBench, which measures the ability to reproduce academic paper experiments; 86.6 on TerminalBench 2.1; 69.0 on Vision2Web; 81.8 on LVBench; and 77.8 on ERQA. Those numbers were also disclosed by Alibaba rather than an outside evaluator.

A model aimed at long-running autonomous work

The article framed the launch against a market where top model developers have become more specialized over the past year. OpenAI’s GPT line has focused on general reasoning and enterprise productivity. Anthropic’s Claude series has leaned into coding and reliable long-context reasoning. Google’s Gemini has pushed multimodal and web-native workflows. Moonshot AI’s Kimi K3 has been building its position around frontier performance and open weights.

Qwen3.8-Max is presented as an attempt to combine those strengths into a single model. The target is not just a chat assistant, but what the report described as an “autonomous employee” capable of working continuously for days.

Alibaba said the model is designed for extended autonomous coding, operating desktop software to handle repetitive business processes, helping research institutions reproduce experiments, and feeding visual input back into a planning loop in manufacturing and logistics settings. In the report’s reading of the launch, the core pitch is not that Qwen3.8-Max chats better, but that it can stay on task for much longer.

2.4 trillion parameters, with about 95 billion active at inference

Qwen3.8-Max uses a mixture-of-experts, or MoE, architecture. Alibaba said the model has 2.4 trillion total parameters, though only about 95 billion are activated during inference. The structure is meant to preserve model capability while reducing computing cost by engaging only a subset of internal “experts” for each request instead of mobilizing the entire network.

The model is built on the Qwen3.5 architecture and expands the context window to 1 million tokens. The article described that as enough to read an entire book in one pass before responding.

Pricing may be the sharper competitive lever

While the benchmark claims drew attention, the report argued that pricing is the part most likely to unsettle rivals. According to the QwenCloud pricing page cited in the piece, overseas pricing is set at $2 per million input tokens and $6 per million output tokens. In mainland China, the listed prices are RMB 12 for input and RMB 36 for output, with implicit cache hits lowering the cost to as little as RMB 1.5.

Using the overseas pricing, the combined input-output cost comes to $8 per million tokens. The article compared that with Claude Opus 5 at a combined $30 per million tokens and GPT-5.6 Sol standard mode at a combined $35. On that basis, Qwen3.8-Max comes in at less than one-third the price of Claude Opus 5 and less than one-quarter the price of GPT-5.6 Sol standard mode.

That matters because autonomous agents consume far more tokens than normal chat interactions. Long workflows involving repeated planning and self-correction can use millions of tokens in a single job. If a company runs hundreds or thousands of agents at once, inference spending can quickly become one of the largest operating costs.

The article also pointed to broader pricing pressure in the sector. OpenAI cut API prices last week for two mid- and lower-tier GPT-5.6 models, Terra and Luna, by 20% and 80%, respectively. In that context, Qwen3.8-Max’s pricing was described as another push deeper into the price war.

Strong in some areas, but not dominant across every benchmark

The report also stressed that Qwen3.8-Max does not top every category. On SWE-Pro, a professional software engineering benchmark, the highest score went to an OpenAI model. In parts of software engineering evaluation and on Agents’ Last Exam, which tests whether a model can autonomously complete long-running tasks, Anthropic’s Opus 4.8 remained in front.

BlockTempo noted that Alibaba’s own comparison table effectively acknowledges this point. The company’s framing is that Qwen3.8-Max offers a balanced mix of performance rather than universal first place across all tasks.

Alibaba also claimed the model can autonomously complete software projects lasting more than 10 days, reproduce research papers involving thousands of lines of code, carry out iterative chip-design optimization, and keep adjusting plans through multimodal feedback. Those demonstrations, however, were also internal and have not been independently verified.

Open weights are promised, but licensing terms are still missing

Alibaba said it plans to release Qwen3.8-Max open weights next week on Hugging Face and ModelScope, while also making the smaller Qwen3.8-27B available. If that happens as described, it would mark the first time a Qwen “Max”-level flagship model has been open for self-hosted deployment rather than being kept mainly behind an API.

Still, the company has only said that it will open the weights. It has not yet disclosed the license terms. The article treated that omission as a central issue rather than a footnote.

If Alibaba uses a permissive license such as Apache 2.0, enterprises could fine-tune the model and integrate it into commercial products more freely. If it adopts a custom structure closer to Moonshot AI’s Kimi K3 licensing, then businesses offering the model externally in a Model-as-a-Service, or MaaS, format may need separate commercial authorization, even if weights are technically open. Until the license is published, companies considering self-deployment still do not know how useful the release will be in practice.

The launch lands in a crowded stretch of the market

Qwen3.8-Max arrives during a particularly busy period for model launches and repositioning. The report said Moonshot AI, OpenAI, and Anthropic have all rolled out new moves within a matter of weeks, with different players emphasizing reasoning, coding, multimodal capability, autonomous agents, or pricing.

Alibaba’s package this time is a combination of benchmark claims that are at or above rival models on selected tests, a price point pushed down to roughly one-quarter of some competitors, a 1 million-token context window, and a promise that open weights are coming next week. But each of those claims still depends on follow-through and outside validation rather than a self-issued scorecard alone.

The article’s conclusion was that benchmark tracks are chosen by vendors, while pricing is where the market casts its own vote. In that framing, the real significance of Qwen3.8-Max is not simply whether it beat a rival on a chart, but whether its pricing forces others to cut further.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.