IBM · Granite 4.2
Mid-size Granite 4.2 dense open-weights reasoning model for self-hosted enterprise tasks.
Leader: Claude Opus 5.5 at 58
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Sonnet 5 at 85.2%
Leader: Fugu at 92.9%
Leader: Claude Opus 5.5 at 1846
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: GLM-5.2 at 92.5%
Leader: Nemotron 3 Ultra 550B A55B at 86.8%
Leader: Kimi K3 at 88.7%
| Result | Score | Reported by | Source |
|---|---|---|---|
| aime-2025 · AIME 2025; thinking | 86.67% | vendor | Hugging Face: ibm-granite/granite-4.2-30b model card ↗ |
| tau2-Bench Banking (AA) · reasoning | 7.6% | independent | Artificial Analysis: Granite 4.2 8B ↗ |
| AA-Omniscience Index · reasoning | -17.2 index | independent | Artificial Analysis: Granite 4.2 8B ↗ |
| SWE-bench Multilingual · thinking | 30.78% | vendor | Hugging Face: ibm-granite/granite-4.2-30b model card ↗ |
| tau3-Bench · thinking | 66.34% | vendor | Hugging Face: ibm-granite/granite-4.2-30b model card ↗ |
| IFBench · prompt-level; thinking | 79.33% | vendor | Hugging Face: ibm-granite/granite-4.2-30b model card ↗ |