Qwen3.8-Flash-Next: Open-Weight Model Beats Claude Opus 4.6 Max
Alibaba introduced Qwen3.8-Flash-Next with weights, and it is already outperforming Claude Opus 4.6 Max in several tests at once.
This is not yet a full-fledged Qwen4, but an early preview of its architecture.
125 billion parameters in the base model plus 51 billion parameters of n-gram representations, with only 6 billion parameters activated per token.
Training cost roughly 9 times less than Qwen3.7-Plus, yet the model proved stronger in programming and office tasks.
Key results:
• SWE-bench Pro — 62.5 vs 53.4 for Claude Opus 4.6 Max
• SWE-bench Multilingual — 81.0
• CoWorkBench — 73.9
• JobBench — 55.7
• Toolathlon — 73.5
• LiveCodeBench — 91.9
• GPQA Diamond — 91.7
• IFBench — 81.3
Native context of 256K tokens, expandable to 1 million via YaRN.
On a 1 million token context, preprocessing speeds up by up to 7.6 times, and generation by up to 4.9 times.
Only 6 billion active parameters per token, yet the level is already comparable to Opus in agentic tasks and programming.
Very soon the model will be available via API at $0.16 per 1 million input tokens and $0.47 per 1 million output tokens.