Mistral · Mistral 3.5
Mistral's dense 128B open-weight flagship that merges instruct, reasoning and coding (replacing Magistral and Devstral 2) for agentic and coding workloads.
Leader: Claude Opus 5.5 at 58
Leader: Claude Opus 5.5 at 66.9%
Leader: Claude Mythos 5.1 at 60.9%
Leader: Claude Fable 5.1 at 91.4%
Leader: Claude Sonnet 5 at 85.2%
Leader: GPT-6 Astra at 96.1%
Leader: Claude Opus 5.5 at 61.4%
Leader: GPT-5.6 Sol at 32.3%
Leader: Kimi K3 at 88.7%
Leader: Claude Opus 5.5 at 87.7%
| Result | Score | Reported by | Source |
|---|---|---|---|
| Terminal-Bench Hard (AA) · high effort | 33.3% | independent | Artificial Analysis: Mistral Medium 3.5 ↗ |
| IFBench · high effort | 68.8% | independent | Artificial Analysis: Mistral Medium 3.5 ↗ |
| tau2-Bench Banking (AA) · high effort | 15.1% | independent | Artificial Analysis: Mistral Medium 3.5 ↗ |
| AA-Omniscience Index · high effort | -36.8 index | independent | Artificial Analysis: Mistral Medium 3.5 ↗ |
| AA Analyst Agent · high effort | 12.5% | independent | Artificial Analysis: Mistral Medium 3.5 ↗ |
| tau3-Bench Telecom | 91.4% | vendor | Mistral Medium 3.5 model card ↗ |