Technology

Qwen3.8-Flash Multimodal MoE Cost-Efficient Model Released With Open Weights; Check Details

Alibaba Qwen has released Qwen3.8-Flash, a multimodal MoE model and early Qwen4 architecture preview. With 125B parameters plus 51B N-gram embeddings and only 6B active per token, it offers strong coding and office performance at roughly one-ninth the training cost of Qwen3.7-Plus. Open weights are available now; the production API arrives soon at low token prices.

Qwen3.8-Flash Multimodal MoE Cost-Efficient Model Released With Open Weights; Check Details
Qwen3.8-Flash (Photo Credits: X/@Alibaba_Qwen)
1
2
3
4
5

Alibaba’s Qwen team has released Qwen3.8-Flash, a multimodal mixture-of-experts model that serves as an early preview of the Qwen4 architecture. The open-weight version, Qwen3.8-Flash-Next, is available now, while the production model will reach the QwenCloud API shortly with competitive pricing.

Alibaba’s Qwen team on Wednesday announced Qwen3.8-Flash, a multimodal MoE model positioned as a cost-efficient early look at the architecture planned for the forthcoming Qwen4 series. The team has made the weights of Qwen3.8-Flash-Next publicly available, allowing developers to examine the new design before the full Qwen4 family arrives. X-Rival Bluesky Now Supports 10-Minute Videos, Faster Uploads.

The model contains 125 billion parameters in its main network plus 51 billion N-gram embedding parameters, yet activates only 6 billion parameters per token. This sparse design underpins its efficiency claims.

Qwen3.8-Flash Price and Availability Details

The production Qwen3.8-Flash will be offered through the QwenCloud API at $0.16 per million input tokens and $0.47 per million output tokens. Weights for the Next variant are already downloadable from major model hubs. The model supports a native context length of 262,144 tokens, which can be extended to 1 million tokens using YaRN.

Training costs are reported at roughly one-ninth those of the earlier Qwen3.7-Plus, while the new model outperforms its predecessor on a range of tasks, particularly coding and office productivity workloads.

Qwen3.8-Flash Architecture and Technical Upgrades

Four main upgrades define the new architecture. Attention uses a hybrid of Gated DeltaNet and Qwen Sparse Attention. Gated Residual expands the residual stream into four branches with dynamic gating. N-gram embeddings expand capacity through local-context lookups that can be stored in host memory. The Muon optimiser has been refined for the new design, with scaling laws adjusted accordingly.

At a 1 million-token context length, the sparse attention kernel delivers substantial speed-ups in both prefill and decode stages compared with previous generations.

Qwen3.8-Flash Benchmark Performance Results

On agentic coding, the model scores 58.7 on DeepSWE 1.1 and 62.5 on SWE-bench Pro. It reaches 73.9 on CoWorkBench for long-horizon office tasks and 84.5 on AndroidWorld for mobile use. Visual mathematics performance stands at 95.7 on MathVision when confidence intervals are included.

In base-model evaluations with only 6 billion active parameters, it leads on eight of fourteen tested benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K, while remaining competitive on the remainder. Why Did Apple Keep Hide My Email on @icloud.com After Announcing a Change?.

The release continues Qwen’s pattern of opening architectural changes early so the community can study them ahead of larger models. Developers can already begin experimenting with the open weights.

Rating:5

TruLY Score 5 – Trustworthy | On a Trust Scale of 0-5 this article has scored 5 on LatestLY. It is verified through official sources (Qwen X Account). The information is thoroughly cross-checked and confirmed. You can confidently share this article with your friends and family, knowing it is trustworthy and reliable.

(The above story first appeared on LatestLY on Aug 26, 2026 08:48 PM IST. For more news and updates on politics, world, sports, entertainment and lifestyle, log on to our website latestly.com).