Qwen3.8-Flash-Next: A Preview of Qwen4 Architecture on NVIDIA GB300 NVL72 for Agentic Coding
Release Overview
Chinese company Alibaba has released the weights of Qwen3.8-Flash-Next as an early preview of the upcoming Qwen4 architecture, allowing developers and researchers to experiment and evaluate before the official launch. This release aims to gather feedback and build an ecosystem around the new architecture.
FAQ
What is Qwen3.8-Flash-Next?
It is a multimodal mixture-of-experts (MoE) model released by Alibaba as a preview of Qwen4 architecture, with 125B main parameters and 51B N-gram embeddings, activating 6B parameters per token, and supporting a 262K-token context.
How does Qwen3.8-Flash-Next compare to other models?
It stands out for its efficiency by activating only 6B parameters per token, reducing computational cost compared to dense models, while offering a long context up to 1M tokens, competing with models like GPT-4 and Claude in coding tasks.
Should MENA development teams adopt this model now?
Yes, teams can start experimenting on NVIDIA GB300 NVL72 to evaluate its performance in agentic coding, especially with long-context support, but it's advisable to wait for the official Qwen4 release for stability.
What are the main use cases for Qwen3.8-Flash-Next?
Agentic coding, long-document analysis, multimodal generation, and building AI agents that require extensive context.
Source: NVIDIA Developer (AI)
AI-assisted content, human-reviewed.