Alibaba · Qwen3.8
Alibaba's flagship hosted Qwen: the managed version of the open Qwen3.8-2.4T-A95B, adding image/video input, non-thinking mode, 1M context and built-in tools. Now on the 0902 post-training snapshot.
Leader: Claude Opus 5.5 at 58
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Opus 5.5 at 1846
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
Leader: Claude Fable 5 at 1506
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| AA Intelligence Index (v4.3.2, reasoning, 0803 launch snapshot (deprecated)) | 40 | independent | Artificial Analysis: Qwen3.8 Max ↗ |
| DeepSWE 1.1 · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 56.6% | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |
| FrontierSWE · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 73.5 score | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |
| IFBench · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 82.8% | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |
| Toolathlon Verified · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 72.5% | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |
| Agents' Last Exam (pass) · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 27% | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |
| MRCR v2 256K (8-needle) · launch version (Qwen3.8-Max column in the Qwen3.8-2.4T-A95B card) | 92.9% | vendor | Qwen3.8-2.4T-A95B model card (Hugging Face) ↗ |