Alibaba open-sources Qwen3.8-Flash-Next as an early look at the Qwen 4 architecture

Alibaba open-sources Qwen3.8-Flash-Next as an early look at the Qwen 4 architecture

N
News Editor
2026-08-26 08:05:56
Alibaba’s Qwen team has released Qwen3.8-Flash-Next on August 26, presenting it as a preview of the Qwen 4 architecture before the full Qwen 4 family arrives. According to the model page on Hugging Face, the release gives developers an early chance to test and prepare for the next-generation architecture rather than wait for the formal rollout. The model is described as an open-source multimodal system built on a mixture-of-experts, or MoE, design. Alibaba said it has about 125 billion total parameters, while activating roughly 6 billion parameters per token. The stated goal is to keep the knowledge scale of a large model while lowering inference costs in practice. Model weights are already available on Hugging Face, and Alibaba said the model is also planned for China’s ModelScope platform. That means developers can access it without relying on a proprietary API. At the same time, Alibaba has not released benchmark results for Qwen3.8-Flash-Next, and the published parameter specifications have not yet been independently verified. Full performance still awaits broader evaluation.

Alibaba’s Qwen team has released an early preview of its next model architecture before the formal debut of Qwen 4. According to the official Hugging Face page, Alibaba published the open-source model Qwen3.8-Flash-Next on August 26 and described it as a preview of the Qwen 4 architecture, giving developers a chance to test the next-generation design before the full Qwen 4 family is introduced.

A 125B MoE model with about 6B parameters activated per token

Based on details released by the team, Qwen3.8-Flash-Next is an open-source multimodal model built with a mixture-of-experts, or MoE, architecture. It has about 125 billion total parameters, while only around 6 billion parameters are activated for each token. Alibaba said that setup is intended to preserve the knowledge capacity of a large model while reducing real inference costs.

The model weights are already listed on Hugging Face, and Alibaba plans to make them available on China’s ModelScope platform as well. Developers can use the model without depending on a proprietary API.

Released as a Qwen 4 architecture preview, with no benchmarks yet

The significance of the release is not limited to the standalone model. It is the first externally released version based on the next-generation Qwen 4 architecture, and its purpose is to let developers get familiar with the system early and prepare for migration.

Alibaba has not published benchmark results for Qwen3.8-Flash-Next. The parameter specifications disclosed so far have also not been independently verified, so its actual performance still requires full evaluation.

Alibaba has kept up a rapid open-source cadence

ABMedia said Alibaba’s recent open-source releases have come in quick succession. In mid-August, the company introduced the smaller Qwen 3.8-27B model, which it said could run on laptops. It has now followed with Flash-Next, putting the Qwen 4 race into view ahead of the full launch.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1800

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.