Anthropic · Claude 5
Anthropic's new default flagship (Sep 2026) for long-running agentic coding and knowledge work, priced below Opus 5.
Leader: Claude Opus 5.5 at 58
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Opus 5.5 at 1846
Leader: Kimi K3 at 84.8%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: GPT-6 Astra at 95%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| SWE-bench Multilingual · max effort, avg of 5 trials | 93.9% | vendor | Claude Opus 5.5 System Card ↗ |
| SWE-bench Multimodal · max effort, avg of 5 trials | 61.4% | vendor | Claude Opus 5.5 System Card ↗ |
| DeepSWE v1.1 · max effort, avg of 5 trials | 74.2% | vendor | Claude Opus 5.5 System Card ↗ |
| FrontierCode v1.1 (Main) · max effort | 54.4% | vendor | Anthropic: Introducing Claude Opus 5.5 ↗ |
| Terminal-Bench-Science 0.1 · max effort | 58.7% | vendor | Anthropic: Introducing Claude Opus 5.5 ↗ |
| AutomationBench · run by Zapier, no fallback models | 40% | vendor | Anthropic: Introducing Claude Opus 5.5 ↗ |
| AA Coding Agent Index v1.5 · Claude Code harness, max effort (DeepSWE 68.4 / Terminal-Bench 4.0 63.1 / SWE-Atlas-QnA 66.4) | 66% | independent | Artificial Analysis Coding Agent Index ↗ |
| ARC-AGI-1 · max effort, semi-private | 97.5% | independent | ARC Prize leaderboard data ↗ |