DeepSeek · DeepSeek V4.1
DeepSeek's current low-cost API default and first natively multimodal (image+text) open-weight MoE, with 552B backbone parameters (8B active in prefill, 16B in decode) and a 1M context.
Leader: Claude Opus 5.5 at 58
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 1846
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: GPT-6 Astra at 96.1%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| Terminal-Bench 3.0 · max effort, DeepSeek Harness (Minimal mode) | 30% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| DeepSWE v1.1 · max effort, mini-SWE harness | 74.2% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| Agents' Last Exam · max effort, official scaffold | 31.8% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| AutomationBench · max effort, official scaffold | 54.8% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| NL2Repo-Bench · max effort | 64 score | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| CyberGym · max effort | 88.1% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| Codeforces rating · max effort | 3471 Elo | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |
| MathArena Apex · max effort | 65.6% | vendor | DeepSeek-V4.1-Flash model card (Hugging Face) ↗ |