OpenAI · GPT-6
OpenAI's mid-tier GPT-6 model for general and coding workloads, replacing GPT-5.6 Sol at lower cost.
Leader: Claude Opus 5.5 at 58
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Opus 5.5 at 1846
Leader: Kimi K3 at 84.8%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| AA-Omniscience Index · max effort | 27.1 index (-100 to 100) | independent | Artificial Analysis: GPT-6 Sol (max) ↗ |
| AutomationBench-AA (partial score) · max effort | 61.6% | independent | Artificial Analysis: GPT-6 Sol (max) ↗ |
| Terminal-Bench Science (AA) · max effort | 30% | independent | Artificial Analysis: GPT-6 Sol (max) ↗ |
| AA Coding Agent Index v1.5 · Codex + GPT-6 Sol (max); DeepSWE 69.0 / SWE-Atlas-QnA 57.5 / TB4 43.4 | 56.7 index | independent | Artificial Analysis Coding Agents ↗ |
| Agents' Last Exam · max | 56.4% | vendor | OpenAI: Introducing GPT-6 Sol and Luna ↗ |
| DeepSWE v1.1 · max | 68.8% | vendor | OpenAI: Introducing GPT-6 Sol and Luna ↗ |
| AutomationBench · xhigh ($0.27/task) | 33.2% | vendor | OpenAI: Introducing GPT-6 Sol and Luna ↗ |