Anthropic · Claude 5
Anthropic's most capable generally available model for demanding reasoning and long-horizon agentic work; successor to Fable 5 at the same token prices.
Leader: Claude Opus 5.5 at 58
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Opus 5.5 at 1846
Leader: Kimi K3 at 84.8%
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: GPT-6 Astra at 95%
Leader: Kimi K3 at 88.7%
Leader: Claude Fable 5 at 1506
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| SWE-bench Multilingual · max effort | 89.1% | vendor | Claude Fable 5.1 / Mythos 5.1 System Card ↗ |
| SWE-bench Multimodal · max effort | 54.7% | vendor | Claude Fable 5.1 / Mythos 5.1 System Card ↗ |
| DeepSWE v1.1 · max effort, avg of 5 trials | 67.4% | vendor | Claude Fable 5.1 / Mythos 5.1 System Card ↗ |
| Terminal-Bench-Science 0.1 · max effort | 52.6% | vendor | Anthropic: Claude Fable 5.1 and Mythos 5.1 ↗ |
| CursorBench 3.2.0 · max effort | 73.4% | vendor | Anthropic: Claude Fable 5.1 and Mythos 5.1 ↗ |
| AutomationBench | 31.4% | vendor | Anthropic: Claude Fable 5.1 and Mythos 5.1 ↗ |
| AA Coding Agent Index v1.5 · Claude Code harness, max effort with fallback (DeepSWE 64.3 / Terminal-Bench 4.0 57.6 / SWE-Atlas-QnA 64.8) | 62.2% | independent | Artificial Analysis Coding Agent Index ↗ |
| ARC-AGI-1 · max effort, semi-private | 97.5% | independent | ARC Prize leaderboard data ↗ |