Moonshot AI · Kimi K3
Moonshot's open-weight flagship: a 2.8T-total / 104B-active multimodal MoE with 1M context, aimed at long-horizon coding and agentic knowledge work. It is the priciest API in this group.
Leader: Claude Opus 5.5 at 58
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Opus 5.5 at 1846
Leader: GPT-6 Astra at 91.5%
Leader: Kimi K3 at 84.2%
Leader: Kimi K3 at 84.8%
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
Leader: Kimi K3 at 90%
Leader: Claude Fable 5 at 1506
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| MMMU-Pro · first of two values reported (81.6 / 83.4); the card does not label the two settings | 81.6% | vendor | Kimi K3 model card (Hugging Face) ↗ |
| DeepSWE · max | 67.5% | vendor | Kimi K3 model card (Hugging Face) ↗ |
| FrontierSWE · max | 81.2 score | vendor | Kimi K3 model card (Hugging Face) ↗ |
| Toolathlon-Verified · max | 76.5% | vendor | Kimi K3 model card (Hugging Face) ↗ |
| Agents' Last Exam · max | 28.3% | vendor | Kimi K3 model card (Hugging Face) ↗ |
| OSWorld 2.0 · max | 58.3% | vendor | Kimi K3 model card (Hugging Face) ↗ |
| Kimi Code Bench 2.0 · max | 72.9% | vendor | Kimi K3 model card (Hugging Face) ↗ |