Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/Mistral
Mistral

Mistral · Devstral

Devstral 2

Open weightsDeprecatedModified MIT

Mistral agentic coding model; API retired 2026-07-31 and replaced by Mistral Medium 3.5 in Vibe.

AA Intelligence
8.6#109 of 142
Price
Self-host
Output speed
—
Context window
256K tokens

Composite indexes

  • Artificial Analysis Intelligence Index#109 of 142
    9independentIntelligence Index v4.3; non-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Opus 5.5 at 58

Coding

  • SciCode#84 of 94
    32.8%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Opus 5.5 at 66.9%

  • Terminal-Bench 4.0#81 of 88
    0%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Mythos 5.1 at 60.9%

  • Terminal-Bench 2.x#75 of 103
    30.3%independentTerminal-Bench 2.1; non-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Fable 5.1 at 91.4%

  • LiveCodeBench#36 of 56
    44.8%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Fugu at 92.9%

Agents and tool use

  • τ²-Bench#61 of 88
    24.9%independenttelecom; non-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: GLM-5.2 at 99.1%

  • GDPval-AA#75 of 103
    531independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Opus 5.5 at 1846

Reasoning and knowledge

  • GPQA Diamond#110 of 142
    59.4%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: GPT-6 Astra at 96.1%

  • Humanity’s Last Exam#145 of 147
    3.6%independentno tools; non-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Claude Opus 5.5 at 61.4%

  • CritPt#109 of 139
    0%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: GPT-5.6 Sol at 32.3%

Vision and long context

  • AA Long Context Reasoning#104 of 139
    32.3%independentnon-reasoningArtificial Analysis: Devstral 2 ↗

    Leader: Kimi K3 at 88.7%

Other published results

ResultScoreReported bySource
aime-2025 · AIME 2025; non-reasoning36.7%independentArtificial Analysis: Devstral 2 ↗
Terminal-Bench Hard (AA) · non-reasoning18.9%independentArtificial Analysis: Devstral 2 ↗
IFBench · non-reasoning38.1%independentArtificial Analysis: Devstral 2 ↗
tau2-Bench Banking (AA) · non-reasoning10.5%independentArtificial Analysis: Devstral 2 ↗
AA-Omniscience Index · non-reasoning-46.7 indexindependentArtificial Analysis: Devstral 2 ↗

Specs

Released
December 9, 2025
Parameters
125B
Size class
100 to 500B
Input
text
Output
text
Context
256K tokens
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • Liquid AILFM2.5-2.6B8.4
    • Sarvam AISarvam 105B8.8
    • GoogleGemma 4 E4B8.9
    • NVIDIANemotron 3 Nano 30B A3B8.9