Thinking Machines Lab · Inkling
Open-weights 276B/12B-active sibling of Inkling (full release after the July preview) that matches or beats Inkling on most benchmarks at roughly a quarter of the cost.
Leader: Claude Opus 5.5 at 58
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Sonnet 5 at 85.2%
Leader: Claude Opus 5.5 at 89.9%
Leader: Claude Opus 5.5 at 1846
Leader: Kimi K3 at 84.2%
Leader: GPT-6 Astra at 91.5%
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: ERNIE 5.1 at 99.6%
Leader: GLM-5.2 at 92.5%
Leader: GPT-6 Astra at 95%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
| Result | Score | Reported by | Source |
|---|---|---|---|
| tau2-Bench Banking (AA) · reasoning | 18.8% | independent | Artificial Analysis: Inkling Small ↗ |
| AA-Omniscience Index · reasoning | -8.9 index | independent | Artificial Analysis: Inkling Small ↗ |
| AA Analyst Agent · reasoning | 27.5% | independent | Artificial Analysis: Inkling Small ↗ |
| Artificial Analysis Intelligence Index v4.1 (as cited by vendor) | 40 index | vendor | Thinking Machines Lab: Introducing Inkling-Small ↗ |
| ARC-AGI-1 · semi-private; XHigh effort; ARC Prize verified | 84% | independent | ARC Prize leaderboard (evaluations data) ↗ |
| IFBench | 82.2% | vendor | Thinking Machines Lab: Introducing Inkling-Small ↗ |
| SimpleQA Verified | 20.6% | vendor | Thinking Machines Lab: Introducing Inkling-Small ↗ |
| Toolathlon Verified | 54.4% | vendor | Thinking Machines Lab: Introducing Inkling-Small ↗ |
| tau3-Bench Banking | 15.5% | vendor | Thinking Machines Lab: Introducing Inkling-Small ↗ |