T/TTokens or Towers
PlannerCatalogScanPricesBenchmarksHarnesses
SITEMAP.XML · checked August 4, 2026
Tokens or Towers prototypeEvery recommendation shows its work.

Benchmarks reference

The evidence, filterable.

A dated snapshot of what providers, model cards, and third-party eval suites report. We link every claim to its source and measure nothing ourselves.

Open the planner →

Scores reflect published claims checked on 2026-08-04. Suites differ in harness, prompting, and tooling, so only compare scores within a single suite.

Filters

Narrow the evidence set.

34 of 34 models

Provider
Weights
Size class
Input types
Output types
  • GPT-5.6 Luna

    OpenAI

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)51

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • ARC-AGI-2 (max)59.5%

      ARC Prize: GPT-5.6 results ↗Checked 2026-08-04

    Fast/affordable tier of the GPT-5.6 family. Parameter count undisclosed.

  • GPT-5.6 Terra

    OpenAI

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)55

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    • Coding Agent Index (max)77

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-05

    Mid GPT-5.6 hosted tier between Luna and Sol. Parameter count undisclosed.

  • GPT-5.6 Sol

    OpenAI

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Agents' Last Exam53.6

      OpenAI: GPT-5.6 Sol preview ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (max)59

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • Coding Agent Index (max, Codex)80

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    Flagship GPT-5.6 model for long-horizon agentic work. Parameter count undisclosed.

  • GPT-5.4 mini

    OpenAI

    closed weights
    Params
    —
    Size
    medium
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)40

      Artificial Analysis: GPT-5.4 mini ↗Checked 2026-08-05

    Compact GPT-5.4 generalist for coding and document work. Parameter count undisclosed.

  • GPT-5.4 nano

    OpenAI

    closed weights
    Params
    —
    Size
    small
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)38

      Artificial Analysis: GPT-5.4 nano vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Lowest-cost GPT-5.4 text tier. Parameter count undisclosed.

  • DeepSeek V4 Flash 0731

    DeepSeek

    open weights
    Params
    284B total · 13B active
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Terminal-Bench 2.182.7

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • DeepSWE54.4

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • Agents' Last Exam25.2

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    MIT open weights; MoE with ~13B active. Official post-training release superseding the preview.

  • DeepSeek V4 Pro

    DeepSeek

    open weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index (max)44

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    Open-weight DeepSeek Pro tier for harder coding and research. Parameter count varies by reporting source.

  • Gemma 4 E4B

    Google

    open weights
    Params
    4B
    Size
    small
    Context
    128K tokens
    in · textin · imagein · audioin · videoout · text
    • GPQA Diamond58.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    • MMMLU76.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    Edge/mobile multimodal utility model from the Gemma 4 family.

  • Gemma 4 26B A4B

    Google

    open weights
    Params
    26B total · 3.8B active
    Size
    medium
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond82.3%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v677.1%

      Gemma 4 model overview ↗Checked 2026-08-04

    MoE local generalist: ~26B total / ~3.8B active. Fits 24–32 GB consumer GPUs when quantized.

  • Gemma 4 31B

    Google

    open weights
    Params
    31B
    Size
    large
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond84.3%

      Gemma 4 model card ↗Checked 2026-08-04

    • MMLU Pro85.2%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v680.0%

      Gemma 4 technical report ↗Checked 2026-08-04

    Dense Gemma 4 flagship for higher-quality local multimodal work.

  • Gemini 3.1 Pro

    Google

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index46

      Artificial Analysis: GPT-5.4 mini vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Hosted Gemini Pro multimodal tier. Parameter count undisclosed.

  • Gemini 3.6 Flash

    Google

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.6 Flash vs Claude Sonnet 5 ↗Checked 2026-08-05

    Fast multimodal Flash generalist on the current Gemini 3.6 line.

  • Gemini 3.5 Flash-Lite

    Google

    closed weights
    Params
    —
    Size
    medium
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index36

      Artificial Analysis: Gemini 3.5 Flash-Lite ↗Checked 2026-08-05

    Low-cost multimodal Flash-Lite route for high-volume document and vision work.

  • Claude Sonnet 5

    Anthropic

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)53

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Premium Claude generalist. Parameter count undisclosed.

  • Claude Haiku 4.5

    Anthropic

    closed weights
    Params
    —
    Size
    medium
    Context
    200K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index30

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Economy Claude route for focused coding and document work.

  • Claude Fable 5

    Anthropic

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index60

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    • Humanity's Last Exam53%

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    Anthropic frontier model; parameter count undisclosed. Some evals fall back to Opus 4.8 under safety guardrails.

  • Claude Opus 4.8

    Anthropic

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Terminal-Bench 2.185.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    • DeepSWE58.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    Prior Anthropic flagship still used as a fallback and comparison baseline.

  • gpt-oss-20b

    OpenAI

    open weights
    Params
    21B total · 3.6B active
    Size
    medium
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)71.5%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)15

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~21B total / 3.6B active). Entry-point coding model for 16 GB-class machines.

  • gpt-oss-120b

    OpenAI

    open weights
    Params
    117B total · 5.1B active
    Size
    large
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)80.1%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • SWE-Bench Verified (high)62.4%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)24

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~117B total / 5.1B active). Capacity-class coding route for 80 GB+ machines.

  • Llama 4 Scout

    Meta

    open weights
    Params
    109B total · 17B active
    Size
    large
    Context
    10M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)74.3%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)57.2%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • MMMU (instruct)69.4%

      Meta Llama 4 announcement ↗Checked 2026-08-04

    Open-weight MoE (17B active / 109B total) with a 10M-token context window.

  • Llama 4 Maverick

    Meta

    open weights
    Params
    400B total · 17B active
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)80.5%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)69.8%

      Meta Llama 4 model card ↗Checked 2026-08-04

    Larger Llama 4 MoE (17B active, 128 experts). Total parameter count commonly reported near 400B.

  • Qwen3.6 27B

    Alibaba

    open weights
    Params
    27B
    Size
    medium
    Context
    262K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index (reasoning)37

      Artificial Analysis model comparison ↗Checked 2026-08-04

    Open-weight mid-size Qwen tier; strong value for local or self-hosted runs.

  • Mistral Small 3.3

    Mistral

    open weights
    Params
    24B
    Size
    medium
    Context
    128K tokens
    in · textin · imageout · text
    • GPQA Diamond (5-shot CoT)46.13%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    • HumanEval Plus (Pass@5)92.90%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    Dense 24B multimodal open weights. Scores cite the published Small 3.2 instruct card for the same size class.

  • Gemini 3.5 Flash

    Google

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.5 Flash ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (launch v4.0)55

      Artificial Analysis: Gemini 3.5 Flash launch ↗Checked 2026-08-04

    Hosted multimodal Flash tier. Index scores shifted between methodology versions; both dated claims are linked.

  • Kimi K3

    Moonshot

    open weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index57

      Santage LLM comparison (citing AA) ↗Checked 2026-08-04

    Open-weight long-horizon coding/reasoning model. Parameter count varies by reporting source.

  • GLM-5

    Z.AI

    open weights
    Params
    744B total · 40B active
    Size
    frontier
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index50

      Artificial Analysis: GLM-5 ↗Checked 2026-08-05

    Open-weight MoE (~744B total / 40B active). Useful open frontier baseline beside Kimi.

  • GLM-5.2

    Z.AI

    open weights
    Params
    744B total · 40B active
    Size
    frontier
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index54

      Artificial Analysis: GLM-5.2 ↗Checked 2026-08-05

    Successor release on the GLM-5 base with stronger agentic coding results.

  • Grok 4.5

    xAI

    closed weights
    Params
    —
    Size
    frontier
    Context
    500K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (high)54

      Artificial Analysis: Grok 4.5 ↗Checked 2026-08-05

    • Artificial Analysis Coding Agent Index (Grok Build)76

      Artificial Analysis: Grok 4.5 launch ↗Checked 2026-08-05

    Current xAI flagship reasoning tier. Parameter count undisclosed.

  • Mistral Large 3

    Mistral

    open weights
    Params
    675B total · 41B active
    Size
    frontier
    Context
    256K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index16

      Artificial Analysis: Mistral Large 3 ↗Checked 2026-08-05

    Open-weight MoE (~675B total / 41B active). Current Mistral Large API tier.

  • Amazon Nova Pro

    Amazon

    closed weights
    Params
    —
    Size
    large
    Context
    300K tokens
    in · textin · imagein · videoout · text
    • MT-Bench (median, Claude 3.7 Sonnet judge)8.5

      AWS: Benchmarking Amazon Nova ↗Checked 2026-08-05

    Bedrock multimodal Pro tier with 300K context. Parameter count undisclosed.

  • Command A

    Cohere

    open weights
    Params
    111B
    Size
    large
    Context
    256K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index8

      Artificial Analysis: Command A ↗Checked 2026-08-05

    Cohere enterprise Command tier. Open weights under a non-commercial research license.

  • MiniMax-M2

    MiniMax

    open weights
    Params
    230B total · 10B active
    Size
    frontier
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index28

      Artificial Analysis: MiniMax-M2 ↗Checked 2026-08-05

    Open-weight MoE (~230B total / 10B active) aimed at agentic coding.

  • Doubao Seed Code

    ByteDance

    closed weights
    Params
    —
    Size
    large
    Context
    256K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index26

      Artificial Analysis: Doubao Seed Code ↗Checked 2026-08-05

    ByteDance Seed coding/reasoning API model. Parameter count undisclosed.

  • ERNIE 4.5 300B A47B

    Baidu

    open weights
    Params
    300B total · 47B active
    Size
    frontier
    Context
    128K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index9

      Artificial Analysis: ERNIE 4.5 300B A47B ↗Checked 2026-08-05

    Open-weight MoE (~300B total / 47B active) from Baidu Ernie.