T/TTokens or Towers
PlannerCatalogScanPricesBenchmarks
US · checked August 4, 2026
Tokens or Towers prototypeEvery recommendation shows its work.

Benchmarks reference

The evidence, filterable.

A dated snapshot of what providers, model cards, and third-party eval suites report. We link every claim to its source and measure nothing ourselves.

Open the planner →

Scores reflect published claims checked on 2026-08-04. Suites differ in harness, prompting, and tooling, so only compare scores within a single suite.

Filters

Narrow the evidence set.

26 of 26 models

Provider
Weights
Size class
Input types
Output types
  • GPT-5.6 Luna

    OpenAI

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)51

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • ARC-AGI-2 (max)59.5%

      ARC Prize: GPT-5.6 results ↗Checked 2026-08-04

    Fast/affordable tier of the GPT-5.6 family. Parameter count undisclosed.

  • GPT-5.6 Terra

    OpenAI

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)55

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    • Coding Agent Index (max)77

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-05

    Mid GPT-5.6 hosted tier between Luna and Sol. Parameter count undisclosed.

  • GPT-5.6 Sol

    OpenAI

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Agents' Last Exam53.6

      OpenAI: GPT-5.6 Sol preview ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (max)59

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • Coding Agent Index (max, Codex)80

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    Flagship GPT-5.6 model for long-horizon agentic work. Parameter count undisclosed.

  • GPT-5.4 mini

    OpenAI

    closed weights
    Params
    —
    Size
    medium
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)40

      Artificial Analysis: GPT-5.4 mini ↗Checked 2026-08-05

    Compact GPT-5.4 generalist for coding and document work. Parameter count undisclosed.

  • GPT-5.4 nano

    OpenAI

    closed weights
    Params
    —
    Size
    small
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)38

      Artificial Analysis: GPT-5.4 nano vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Lowest-cost GPT-5.4 text tier. Parameter count undisclosed.

  • DeepSeek V4 Flash 0731

    DeepSeek

    open weights
    Params
    284B total · 13B active
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Terminal-Bench 2.182.7

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • DeepSWE54.4

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • Agents' Last Exam25.2

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    MIT open weights; MoE with ~13B active. Official post-training release superseding the preview.

  • DeepSeek V4 Pro

    DeepSeek

    open weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index (max)44

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    Open-weight DeepSeek Pro tier for harder coding and research. Parameter count varies by reporting source.

  • Gemma 4 E4B

    Google

    open weights
    Params
    4B
    Size
    small
    Context
    128K tokens
    in · textin · imagein · audioin · videoout · text
    • GPQA Diamond58.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    • MMMLU76.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    Edge/mobile multimodal utility model from the Gemma 4 family.

  • Gemma 4 26B A4B

    Google

    open weights
    Params
    26B total · 3.8B active
    Size
    medium
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond82.3%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v677.1%

      Gemma 4 model overview ↗Checked 2026-08-04

    MoE local generalist: ~26B total / ~3.8B active. Fits 24–32 GB consumer GPUs when quantized.

  • Gemma 4 31B

    Google

    open weights
    Params
    31B
    Size
    large
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond84.3%

      Gemma 4 model card ↗Checked 2026-08-04

    • MMLU Pro85.2%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v680.0%

      Gemma 4 technical report ↗Checked 2026-08-04

    Dense Gemma 4 flagship for higher-quality local multimodal work.

  • Gemini 3.1 Pro

    Google

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index46

      Artificial Analysis: GPT-5.4 mini vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Hosted Gemini Pro multimodal tier. Parameter count undisclosed.

  • Gemini 3.6 Flash

    Google

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.6 Flash vs Claude Sonnet 5 ↗Checked 2026-08-05

    Fast multimodal Flash generalist on the current Gemini 3.6 line.

  • Gemini 3.5 Flash-Lite

    Google

    closed weights
    Params
    —
    Size
    medium
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index36

      Artificial Analysis: Gemini 3.5 Flash-Lite ↗Checked 2026-08-05

    Low-cost multimodal Flash-Lite route for high-volume document and vision work.

  • Claude Sonnet 5

    Anthropic

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)53

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Premium Claude generalist. Parameter count undisclosed.

  • Claude Haiku 4.5

    Anthropic

    closed weights
    Params
    —
    Size
    medium
    Context
    200K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index30

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Economy Claude route for focused coding and document work.

  • Claude Fable 5

    Anthropic

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index60

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    • Humanity's Last Exam53%

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    Anthropic frontier model; parameter count undisclosed. Some evals fall back to Opus 4.8 under safety guardrails.

  • Claude Opus 4.8

    Anthropic

    closed weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • Terminal-Bench 2.185.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    • DeepSWE58.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    Prior Anthropic flagship still used as a fallback and comparison baseline.

  • gpt-oss-20b

    OpenAI

    open weights
    Params
    21B total · 3.6B active
    Size
    medium
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)71.5%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)15

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~21B total / 3.6B active). Entry-point coding model for 16 GB-class machines.

  • gpt-oss-120b

    OpenAI

    open weights
    Params
    117B total · 5.1B active
    Size
    large
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)80.1%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • SWE-Bench Verified (high)62.4%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)24

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~117B total / 5.1B active). Capacity-class coding route for 80 GB+ machines.

  • Llama 4 Scout

    Meta

    open weights
    Params
    109B total · 17B active
    Size
    large
    Context
    10M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)74.3%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)57.2%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • MMMU (instruct)69.4%

      Meta Llama 4 announcement ↗Checked 2026-08-04

    Open-weight MoE (17B active / 109B total) with a 10M-token context window.

  • Llama 4 Maverick

    Meta

    open weights
    Params
    400B total · 17B active
    Size
    frontier
    Context
    1M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)80.5%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)69.8%

      Meta Llama 4 model card ↗Checked 2026-08-04

    Larger Llama 4 MoE (17B active, 128 experts). Total parameter count commonly reported near 400B.

  • Qwen3.6 27B

    Alibaba

    open weights
    Params
    27B
    Size
    medium
    Context
    262K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index (reasoning)37

      Artificial Analysis model comparison ↗Checked 2026-08-04

    Open-weight mid-size Qwen tier; strong value for local or self-hosted runs.

  • Mistral Small 3.3

    Mistral

    open weights
    Params
    24B
    Size
    medium
    Context
    128K tokens
    in · textin · imageout · text
    • GPQA Diamond (5-shot CoT)46.13%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    • HumanEval Plus (Pass@5)92.90%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    Dense 24B multimodal open weights. Scores cite the published Small 3.2 instruct card for the same size class.

  • Gemini 3.5 Flash

    Google

    closed weights
    Params
    —
    Size
    large
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.5 Flash ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (launch v4.0)55

      Artificial Analysis: Gemini 3.5 Flash launch ↗Checked 2026-08-04

    Hosted multimodal Flash tier. Index scores shifted between methodology versions; both dated claims are linked.

  • Kimi K3

    Moonshot

    open weights
    Params
    —
    Size
    frontier
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index57

      Santage LLM comparison (citing AA) ↗Checked 2026-08-04

    Open-weight long-horizon coding/reasoning model. Parameter count varies by reporting source.

  • GLM-5

    Z.AI

    open weights
    Params
    744B total · 40B active
    Size
    frontier
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index50

      Artificial Analysis: GLM-5 ↗Checked 2026-08-05

    Open-weight MoE (~744B total / 40B active). Useful open frontier baseline beside Kimi.