Alibaba Previews Next-Generation LLM "Qwen3.8-Flash-Next"
Alibaba's Qwen team has released a preview of "Qwen3.8-Flash-Next," an early demonstration of the next-generation Qwen4 architecture. Adopting a Mixture of Experts (MoE) structure with 1.25 trillion total parameters but only 6 billion active during inference, the model achieves performance surpassing DeepSeek-V4-Flash and Claude Opus 4.6 on some benchmarks while requiring approximately one-ninth the training cost of comparable competitors.

Alibaba's Qwen team has released a preview version of a language model called "Qwen3.8-Flash-Next" that adopts a new architecture. This model is positioned as an early demonstration of the Qwen4-generation architecture and is reported to achieve high performance while significantly reducing training costs.
The mechanism adopted by this model, called "MoE (Mixture of Experts)," involves maintaining a large number of parameters—an indicator of a model's learned knowledge—while using only a portion during actual inference. Specifically, of the total 1.25 trillion parameters, only 6 billion are activated each time a single token (a unit of words) is processed. Since output can be generated by running less than 5% of the total parameters, computational costs can be substantially reduced.
According to the Qwen team, the training cost was approximately one-ninth that of comparable competitors at the same scale. Nevertheless, on benchmarks for coding and office-related tasks, the model surpassed larger models such as DeepSeek's DeepSeek-V4-Flash and Anthropic's Claude Opus 4.6.
This combination of "low cost and high performance" merits attention in the context of price competition for AI models. It can be seen as potentially creating new pressure on major providers like OpenAI and Anthropic in terms of cost efficiency. In the AI industry, while model performance continues to improve, API usage fees show a downward trend, and this announcement may further accelerate that trajectory.
Additionally, this release has the character of an "advance preview" before Alibaba launches the full Qwen4 generation. It is a disclosure at the stage of presenting the direction of the architecture to the outside world and can be viewed as a technical foundation for future official releases. The utilization of MoE structures is a technique that Chinese AI developers, including DeepSeek, have actively promoted over recent years, and Alibaba is advancing this approach further.
If the coexistence of cost efficiency and model performance is demonstrated, companies and developers utilizing AI through APIs may see expanded access to more affordable services. However, the current release is merely a preview, and details regarding actual product deployment and pricing will need to await future official announcements. How Qwen4's full scope emerges will be the next point of attention.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.