xAI · Grok 4
xAI's current flagship Grok model on a new larger base, with low to xhigh reasoning effort, aimed at coding and agentic work.
Leader: Claude Opus 5.5 at 58
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Opus 5.5 at 1846
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 1818
| Result | Score | Reported by | Source |
|---|---|---|---|
| AA-Omniscience Index · xhigh | 32 index (-100 to 100) | independent | Artificial Analysis: Grok 4.7 (xhigh) ↗ |
| AutomationBench-AA (partial score) · xhigh | 65.6% | independent | Artificial Analysis: Grok 4.7 (xhigh) ↗ |
| AA Coding Agent Index v1.5 · Grok Build + Grok 4.7 (xhigh); DeepSWE 72.6 / SWE-Atlas-QnA 62.9 / TB4 33.3 | 56.3 index | independent | Artificial Analysis Coding Agents ↗ |
| AA Intelligence Index (high) · AA Intelligence Index v4.3.2, high | 46.3 index | independent | Artificial Analysis: Grok 4.7 (high) ↗ |
| CursorBench 4.0 · xHigh | 46.3% | vendor | xAI: Grok 4.7 ↗ |
| DeepSWE v1.1 · high | 71% | vendor | xAI: Grok 4.7 ↗ |
| Terminal-Bench 4.0 (vendor) · xHigh | 37.6% | vendor | xAI: Grok 4.7 ↗ |
| EEBench · xHigh | 64% | vendor | xAI: Grok 4.7 ↗ |
| HealthBench Professional · xHigh | 56.7% | vendor | xAI: Grok 4.7 ↗ |
| Harvey Legal Agent Benchmark · xHigh | 19.6% | vendor | xAI: Grok 4.7 ↗ |