Alibaba has officially released and open-sourced its latest Qwen model, the Qwen3.8-Flash, marking a significant leap in AI efficiency and capability. This new model leverages a cutting-edge next-generation architecture, requiring only 6 billion (6B) activated parameters out of its massive 100 billion total parameter count to achieve frontier-level performance that surpasses Claude Opus4.6, setting a new global benchmark for model efficiency.
Thanks to groundbreaking innovations in both architecture and training methodology, the training cost for Qwen3.8-Flash has plummeted by nearly 90% compared to its predecessor, Qwen3.7-Plus. Furthermore, inference costs have been dramatically reduced, with pricing as low as 1 RMB per million input tokens and 3 RMB per million output tokens, making it the most affordable option at just one-third the price of DeepSeek-V4-Flash. Starting tonight, Qwen3.8-Flash will debut on the Qwen Office platform, with developers and enterprises able to access the new model's API services through the Qwen AI platform.
Qwen3.8-Flash is a multimodal Mixture-of-Experts (MoE) model built on a next-generation (Next) architecture. Despite having 125 billion Transformer parameters, it activates only 6 billion, delivering performance that exceeds Opus4.6 and approaches that of Opus4.8. Even in its pre-training phase, the Qwen3.8-Flash Base model outperforms the Qwen3.7-Plus Base model—which is three times its size—in foundational capabilities such as general knowledge (SuperGPQA), mathematical reasoning (GSM8K), and programming (SWEBench-Pretrain).
Following post-training, Qwen3.8-Flash's performance undergoes a substantial leap: it leads Opus4.6 by an impressive 9.1 points in the SWE-bench Pro evaluation, which tests agentic coding abilities. In other agentic tasks, including the long-horizon professional benchmark CoWorkBench and the realistic tool-use evaluation Toolathlon Verified, it also surpasses DeepSeek-V4-Flash.
The technological advancements in the Qwen3.8 series have resulted in substantial cost reductions for both training and inference. Qwen3.8-Flash achieves performance comparable to the previous generation Qwen3.7-Plus, but it was trained using only one-ninth of the resources, translating to a 90% decrease in training costs. For inference, Qwen3.8-Flash offers input tokens at just 1 RMB per million and output tokens at 3 RMB per million—a mere 3% of the cost of Claude Opus4.6, two-thirds the off-peak price, and one-third the peak-time price of DeepSeek-V4-Flash.