Alibaba’s Qwen team has released an early preview of its next model architecture before the formal debut of Qwen 4. According to the official Hugging Face page, Alibaba published the open-source model Qwen3.8-Flash-Next on August 26 and described it as a preview of the Qwen 4 architecture, giving developers a chance to test the next-generation design before the full Qwen 4 family is introduced.
A 125B MoE model with about 6B parameters activated per token
Based on details released by the team, Qwen3.8-Flash-Next is an open-source multimodal model built with a mixture-of-experts, or MoE, architecture. It has about 125 billion total parameters, while only around 6 billion parameters are activated for each token. Alibaba said that setup is intended to preserve the knowledge capacity of a large model while reducing real inference costs.
The model weights are already listed on Hugging Face, and Alibaba plans to make them available on China’s ModelScope platform as well. Developers can use the model without depending on a proprietary API.
Released as a Qwen 4 architecture preview, with no benchmarks yet
The significance of the release is not limited to the standalone model. It is the first externally released version based on the next-generation Qwen 4 architecture, and its purpose is to let developers get familiar with the system early and prepare for migration.
Alibaba has not published benchmark results for Qwen3.8-Flash-Next. The parameter specifications disclosed so far have also not been independently verified, so its actual performance still requires full evaluation.
Alibaba has kept up a rapid open-source cadence
ABMedia said Alibaba’s recent open-source releases have come in quick succession. In mid-August, the company introduced the smaller Qwen 3.8-27B model, which it said could run on laptops. It has now followed with Flash-Next, putting the Qwen 4 race into view ahead of the full launch.

