Skip to contentT/TTokens or Towers
01Planner02Shortlist03Scan04Prices05Benchmarks06Findings07Harnesses08Hosts09Enterprise
US · checked August 4, 2026
US · checked August 4, 2026
Tokens or Towers prototypeHow we workEvery recommendation shows its work.

Benchmarks reference

The evidence, filterable.

A dated snapshot of provider claims and independent evaluations. Every claim links to its source. All measurements come from their publishers.

Open the planner →

Size buckets follow total parameter count, from under 10B up to the trillion-scale ranges, and each entry also lists its active parameters where published. Closed models with unpublished counts appear as Undisclosed. Scores reflect claims checked through 2026-08-05. Eval setups vary, so compare scores only within the same suite.

Filters

Narrow the evidence set.

67 of 67 models

Provider
Weights
Size class
Input types
Output types
  • GPT-5.6 Luna

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)51

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • ARC-AGI-2 (max)59.5%

      ARC Prize: GPT-5.6 results ↗Checked 2026-08-04

    Fast/affordable tier of the GPT-5.6 family. Parameter count undisclosed.

  • GPT-5.6 Terra

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)55

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    • Coding Agent Index (max)77

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-05

    Mid GPT-5.6 hosted tier between Luna and Sol. Parameter count undisclosed.

  • GPT-5.6 Sol

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Agents' Last Exam53.6

      OpenAI: GPT-5.6 Sol preview ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (max)59

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    • Coding Agent Index (max, Codex)80

      Artificial Analysis: GPT-5.6 ↗Checked 2026-08-04

    Flagship GPT-5.6 model for long-horizon agentic work. Parameter count undisclosed.

  • GPT-5.4 mini

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)40

      Artificial Analysis: GPT-5.4 mini ↗Checked 2026-08-05

    Compact GPT-5.4 generalist for coding and document work. Parameter count undisclosed.

  • GPT-5.4 nano

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    400K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (xhigh)38

      Artificial Analysis: GPT-5.4 nano vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Lowest-cost GPT-5.4 text tier. Parameter count undisclosed.

  • DeepSeek V4 Flash 0731

    DeepSeekDeepSeek

    open weights100 to 500B
    Params
    284B total · 13B active
    Context
    1M tokens
    in · textout · text
    • Terminal-Bench 2.182.7

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • DeepSWE54.4

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    • Agents' Last Exam25.2

      DeepSeek V4 Flash 0731 model card ↗Checked 2026-08-04

    MIT open weights; MoE with ~13B active. Official post-training release superseding the preview.

  • DeepSeek V4 Pro

    DeepSeekDeepSeek

    open weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index (max)44

      Artificial Analysis: GPT-5.6 Terra vs DeepSeek V4 Pro ↗Checked 2026-08-05

    Open-weight DeepSeek Pro tier for harder coding and research. Parameter count varies by reporting source.

  • Gemma 4 E4B

    GoogleGoogle

    open weightsUnder 10B
    Params
    4B
    Context
    128K tokens
    in · textin · imagein · audioin · videoout · text
    • GPQA Diamond58.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    • MMMLU76.6%

      Gemma 4 technical report ↗Checked 2026-08-04

    Edge/mobile multimodal utility model from the Gemma 4 family.

  • Gemma 4 26B A4B

    GoogleGoogle

    open weights10 to 40B
    Params
    26B total · 3.8B active
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond82.3%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v677.1%

      Gemma 4 model overview ↗Checked 2026-08-04

    MoE local generalist: ~26B total / ~3.8B active. Fits 24–32 GB consumer GPUs when quantized.

  • Gemma 4 31B

    GoogleGoogle

    open weights10 to 40B
    Params
    31B
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA Diamond84.3%

      Gemma 4 model card ↗Checked 2026-08-04

    • MMLU Pro85.2%

      Gemma 4 technical report ↗Checked 2026-08-04

    • LiveCodeBench v680.0%

      Gemma 4 technical report ↗Checked 2026-08-04

    Dense Gemma 4 flagship for higher-quality local multimodal work.

  • Gemini 3.1 Pro

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index46

      Artificial Analysis: GPT-5.4 mini vs Gemini 3.1 Pro ↗Checked 2026-08-05

    Hosted Gemini Pro multimodal tier. Parameter count undisclosed.

  • Gemini 3.6 Flash

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.6 Flash vs Claude Sonnet 5 ↗Checked 2026-08-05

    Fast multimodal Flash generalist on the current Gemini 3.6 line.

  • Gemini 3.5 Flash-Lite

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index36

      Artificial Analysis: Gemini 3.5 Flash-Lite ↗Checked 2026-08-05

    Low-cost multimodal Flash-Lite route for high-volume document and vision work.

  • Claude Sonnet 5

    AnthropicAnthropic

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)53

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Premium Claude generalist. Parameter count undisclosed.

  • Claude Haiku 4.5

    AnthropicAnthropic

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    200K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index30

      Artificial Analysis: Anthropic models ↗Checked 2026-08-05

    Economy Claude route for focused coding and document work.

  • Claude Fable 5

    AnthropicAnthropic

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index60

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    • Humanity's Last Exam53%

      Artificial Analysis: Claude Fable 5 ↗Checked 2026-08-04

    Anthropic frontier model; parameter count undisclosed. Some evals fall back to Opus 4.8 under safety guardrails.

  • Claude Opus 4.8

    AnthropicAnthropic

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Terminal-Bench 2.185.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    • DeepSWE58.0

      DeepSeek V4 Flash 0731 comparison table ↗Checked 2026-08-04

    Prior Anthropic flagship still used as a fallback and comparison baseline.

  • Claude Opus 5

    AnthropicAnthropic

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (max)61

      Artificial Analysis: Claude Opus 5 ↗Checked 2026-08-05

    Current Claude Opus API tier. Parameter count undisclosed.

  • gpt-oss-20b

    OpenAIOpenAI

    open weights10 to 40B
    Params
    21B total · 3.6B active
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)71.5%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)15

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~21B total / 3.6B active). Entry-point coding model for 16 GB-class machines.

  • gpt-oss-120b

    OpenAIOpenAI

    open weights100 to 500B
    Params
    117B total · 5.1B active
    Context
    128K tokens
    in · textout · text
    • GPQA Diamond (high, no tools)80.1%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • SWE-Bench Verified (high)62.4%

      OpenAI gpt-oss capability findings ↗Checked 2026-08-05

    • Artificial Analysis Intelligence Index (high)24

      Artificial Analysis: gpt-oss-120b vs gpt-oss-20b ↗Checked 2026-08-05

    Open-weight MoE (~117B total / 5.1B active). Capacity-class coding route for 80 GB+ machines.

  • Llama 4 Scout

    MetaMeta

    open weights100 to 500B
    Params
    109B total · 17B active
    Context
    10M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)74.3%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)57.2%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • MMMU (instruct)69.4%

      Meta Llama 4 announcement ↗Checked 2026-08-04

    Open-weight MoE (17B active / 109B total) with a 10M-token context window.

  • Llama 4 Maverick

    MetaMeta

    open weights100 to 500B
    Params
    400B total · 17B active
    Context
    1M tokens
    in · textin · imageout · text
    • MMLU Pro (instruct)80.5%

      Meta Llama 4 model card ↗Checked 2026-08-04

    • GPQA Diamond (instruct)69.8%

      Meta Llama 4 model card ↗Checked 2026-08-04

    Larger Llama 4 MoE (17B active, 128 experts). Total parameter count commonly reported near 400B.

  • Qwen3.6 27B

    AlibabaAlibaba

    open weights10 to 40B
    Params
    27B
    Context
    262K tokens
    in · textin · imagein · videoout · text
    • Artificial Analysis Intelligence Index (reasoning)37

      Artificial Analysis model comparison ↗Checked 2026-08-04

    Open-weight mid-size Qwen tier; strong value for local or self-hosted runs.

  • Qwen 3.7 Max

    AlibabaAlibaba

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index46

      Artificial Analysis: Qwen3.7 Max ↗Checked 2026-08-05

    Hosted Qwen Max tier. Parameter count undisclosed.

  • Qwen 3.7 Plus

    AlibabaAlibaba

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imagein · videoout · text
    • Artificial Analysis Intelligence Index39

      Artificial Analysis: Qwen3.7 Plus ↗Checked 2026-08-05

    Hosted multimodal Qwen Plus tier. Parameter count undisclosed.

  • Qwen 3.6 Flash

    AlibabaAlibaba

    open weights10 to 40B
    Params
    35B total · 3B active
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • GPQA86.0

      Qwen: Qwen3.6-35B-A3B ↗Checked 2026-08-05

    Hosted alias for the open Qwen3.6 35B-A3B checkpoint.

  • Qwen3.6-35B-A3B

    AlibabaAlibaba

    open weights10 to 40B
    Params
    35B total · 3B active
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • SWE-bench Verified73.4

      Qwen: Qwen3.6-35B-A3B ↗Checked 2026-08-05

    • VideoMME without subtitles82.5

      Qwen: Qwen3.6-35B-A3B ↗Checked 2026-08-05

    Open MoE release behind the Qwen3.6 Flash API tier.

  • Qwen3.8-Max

    AlibabaAlibaba

    closed weights2 to 4T
    Params
    2.4T total · 95B active
    Context
    Not published
    in · textin · imageout · text
    • Release announcementAugust 2, 2026

      Qwen: Qwen3.8-Max release ↗Checked 2026-08-05

    • Family overviewQwen3.8

      OpenLM: Qwen3.8 ↗Checked 2026-08-05

    The first Max-class Qwen with announced open weights, scheduled for release the week after the August 2 launch. 2.4T total with 95B active. This entry flips to open the day the weights land.

  • Qwen3.7-235B-A22B

    AlibabaAlibaba

    open weights100 to 500B
    Params
    235B total · 22B active
    Context
    262K tokens
    in · textout · text
    • LiveCodeBench v679.4%

      Qwen3.7 model card ↗Checked 2026-08-05

    Open-weight MoE flagship of the Qwen3.7 line.

  • Mistral Small 3.3

    MistralMistral

    open weights10 to 40B
    Params
    24B
    Context
    128K tokens
    in · textin · imageout · text
    • GPQA Diamond (5-shot CoT)46.13%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    • HumanEval Plus (Pass@5)92.90%

      Mistral Small 3.2 model card ↗Checked 2026-08-05

    Dense 24B multimodal open weights. Scores cite the published Small 3.2 instruct card for the same size class.

  • Gemini 3.5 Flash

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imagein · audioin · videoout · text
    • Artificial Analysis Intelligence Index (high)50

      Artificial Analysis: Gemini 3.5 Flash ↗Checked 2026-08-04

    • Artificial Analysis Intelligence Index (launch v4.0)55

      Artificial Analysis: Gemini 3.5 Flash launch ↗Checked 2026-08-04

    Hosted multimodal Flash tier. Index scores shifted between methodology versions; both dated claims are linked.

  • Kimi K3

    Moonshot AIMoonshot AI

    open weights2 to 4T
    Params
    2.8T
    Context
    1M tokens
    in · textout · text
    • Artificial Analysis Intelligence Index57

      Santage LLM comparison (citing AA) ↗Checked 2026-08-04

    Open-weight long-horizon coding/reasoning model, commonly reported near 2.8T total parameters.

  • Kimi K2.7 Code

    Moonshot AIMoonshot AI

    open weights1 to 2T
    Params
    1T total · 32B active
    Context
    256K tokens
    in · textin · imagein · videoout · text
    • Kimi Code Bench v262.0

      Moonshot AI: Kimi K2.7 Code model card ↗Checked 2026-08-05

    • MCP Atlas76.0

      Moonshot AI: Kimi K2.7 Code model card ↗Checked 2026-08-05

    Open coding model with 1T total parameters and 32B active per token.

  • GLM-5

    Z.aiZ.ai

    open weights500B to 1T
    Params
    744B total · 40B active
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index50

      Artificial Analysis: GLM-5 ↗Checked 2026-08-05

    Open-weight MoE (~744B total / 40B active). Useful open frontier baseline beside Kimi.

  • GLM-5.2

    Z.aiZ.ai

    open weights500B to 1T
    Params
    744B total · 40B active
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index54

      Artificial Analysis: GLM-5.2 ↗Checked 2026-08-05

    Successor release on the GLM-5 base with stronger agentic coding results.

  • Grok 4.3

    xAIxAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (high)38

      Artificial Analysis: Grok 4.3 ↗Checked 2026-08-05

    Prior Grok API tier retained for comparison. Parameter count undisclosed.

  • Grok Composer 2.5

    xAIxAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Not published
    in · textout · text
    • Release capabilityLong-running coding

      xAI: Composer 2.5 ↗Checked 2026-08-05

    Coding model available through Grok Build. Parameter count undisclosed.

  • Grok 4.5

    xAIxAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    500K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index (high)54

      Artificial Analysis: Grok 4.5 ↗Checked 2026-08-05

    • Artificial Analysis Coding Agent Index (Grok Build)76

      Artificial Analysis: Grok 4.5 launch ↗Checked 2026-08-05

    Current xAI flagship reasoning tier. Parameter count undisclosed.

  • Mistral Large 3

    MistralMistral

    open weights500B to 1T
    Params
    675B total · 41B active
    Context
    256K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index16

      Artificial Analysis: Mistral Large 3 ↗Checked 2026-08-05

    Open-weight MoE (~675B total / 41B active). Current Mistral Large API tier.

  • Amazon Nova Pro

    AmazonAmazon

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    300K tokens
    in · textin · imagein · videoout · text
    • MT-Bench (median, Claude 3.7 Sonnet judge)8.5

      AWS: Benchmarking Amazon Nova ↗Checked 2026-08-05

    Bedrock multimodal Pro tier with 300K context. Parameter count undisclosed.

  • Command A

    CohereCohere

    open weights100 to 500B
    Params
    111B
    Context
    256K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index8

      Artificial Analysis: Command A ↗Checked 2026-08-05

    Cohere enterprise Command tier. Open weights under a non-commercial research license.

  • MiniMax-M2

    MiniMaxMiniMax

    open weights100 to 500B
    Params
    230B total · 10B active
    Context
    200K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index28

      Artificial Analysis: MiniMax-M2 ↗Checked 2026-08-05

    Open-weight MoE (~230B total / 10B active) aimed at agentic coding.

  • Doubao Seed Code

    ByteDanceByteDance

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    256K tokens
    in · textin · imageout · text
    • Artificial Analysis Intelligence Index26

      Artificial Analysis: Doubao Seed Code ↗Checked 2026-08-05

    ByteDance Seed coding/reasoning API model. Parameter count undisclosed.

  • ERNIE 4.5 300B A47B

    Baidu

    open weights100 to 500B
    Params
    300B total · 47B active
    Context
    128K tokens
    in · textout · text
    • Artificial Analysis Intelligence Index9

      Artificial Analysis: ERNIE 4.5 300B A47B ↗Checked 2026-08-05

    Open-weight MoE (~300B total / 47B active) from Baidu Ernie.

  • Inkling

    Thinking Machines Lab

    open weights500B to 1T
    Params
    975B total · 41B active
    Context
    1M tokens
    in · textin · imagein · audioout · text
    • Humanity's Last Exam (text only)29.7%

      Thinking Machines Lab: Inkling ↗Checked 2026-08-05

    • SWE-bench Verified77.6%

      Thinking Machines Lab: Inkling ↗Checked 2026-08-05

    Open MoE model with 975B total parameters and 41B active per token.

  • Inkling-Small Preview

    Thinking Machines Lab

    closed weights100 to 500B
    Params
    276B total · 12B active
    Context
    256K tokens
    in · textin · imagein · audioout · text
    • GPQA Diamond88.3%

      Thinking Machines Lab: Inkling-Small preview ↗Checked 2026-08-05

    Preview checkpoint with 276B total parameters and 12B active. Full weights are pending.

  • Nemotron 3 Nano Omni

    NVIDIA

    open weights10 to 40B
    Params
    30B total · 3B active
    Context
    256K tokens
    in · textin · imagein · audioin · videoout · text
    • VideoMME without subtitles (reasoning)72.2

      NVIDIA: Nemotron 3 Nano Omni report ↗Checked 2026-08-05

    Open omni model with 30B total parameters and 3B active per token.

  • Nemotron 3 Super

    NVIDIA

    open weights100 to 500B
    Params
    120B total · 12B active
    Context
    1M tokens
    in · textout · text
    • HMMT February 202595.4

      NVIDIA: Nemotron 3 Super report ↗Checked 2026-08-05

    Open agentic MoE with 120B total parameters and 12B active per token.

  • Nemotron 3 Ultra

    NVIDIA

    open weights500B to 1T
    Params
    550B total · 55B active
    Context
    1M tokens
    in · textout · text
    • Terminal-Bench 2.156.4

      NVIDIA: Nemotron 3 Ultra report ↗Checked 2026-08-05

    Open Nemotron flagship with 550B total parameters and 55B active per token.

  • Mistral Small 4

    MistralMistral

    open weights100 to 500B
    Params
    119B total · 6B active
    Context
    256K tokens
    in · textin · imageout · text
    • AA-LCR0.72

      Mistral: Mistral Small 4 ↗Checked 2026-08-05

    Apache 2.0 MoE with 119B total parameters and 6B active per token.

  • Phi-4

    Microsoft

    open weights10 to 40B
    Params
    14B
    Context
    16K tokens
    in · textout · text
    • GPQA56.1

      Microsoft: Phi-4 model card ↗Checked 2026-08-05

    Dense 14B open model focused on compact reasoning workloads.

  • Phi-4 Mini Instruct

    Microsoft

    open weightsUnder 10B
    Params
    3.8B
    Context
    128K tokens
    in · textout · text
    • MMLU-Pro (zero-shot CoT)52.8

      Microsoft: Phi-4 Mini model card ↗Checked 2026-08-05

    Dense 3.8B open model for local text inference.

  • Phi-4 Multimodal Instruct

    Microsoft

    open weightsUnder 10B
    Params
    5.6B
    Context
    128K tokens
    in · textin · imagein · audioout · text
    • MathVista testmini62.4

      Microsoft: Phi-4 Multimodal model card ↗Checked 2026-08-05

    Open 5.6B model for multimodal understanding.

  • MiniMax-M3

    MiniMaxMiniMax

    open weights100 to 500B
    Params
    428B total · 23B active
    Context
    1M tokens
    in · textin · imagein · videoout · text
    • Artificial Analysis Intelligence Index44

      Artificial Analysis: MiniMax-M3 ↗Checked 2026-08-05

    • BrowseComp83.5

      MiniMax: M3 ↗Checked 2026-08-05

    Open multimodal MoE with 428B total parameters and 23B active per token.

  • Fugu Ultra

    Sakana AISakana AI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • SWE-bench Pro73.7%

      Sakana: Fugu release ↗Checked 2026-08-05

    • Terminal-Bench 2.182.1%

      Sakana: Fugu release ↗Checked 2026-08-05

    • GPQA Diamond95.5%

      Sakana: Fugu release ↗Checked 2026-08-05

    Orchestration model that routes hard tasks to a pool of expert models. Scores are Sakana-reported and independent verification is pending.

  • Fugu

    Sakana AISakana AI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    1M tokens
    in · textin · imageout · text
    • SWE-bench Pro59.0%

      Sakana: Fugu release ↗Checked 2026-08-05

    • GPQA Diamond95.5%

      Sakana: Fugu release ↗Checked 2026-08-05

    Lower-latency orchestration tier that bills at the routed model rate. Scores are Sakana-reported and independent verification is pending.

  • Seedance 2.0

    ByteDanceByteDance

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 15s video
    in · textin · imagein · audioin · videoout · video
    • Artificial Analysis T2V with audio1226 Elo

      Artificial Analysis: video leaderboard ↗Checked 2026-08-05

    Native multimodal video generation with synchronized audio.

  • Seedance 2.5

    ByteDanceByteDance

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Enterprise beta
    in · textin · imagein · audioin · videoout · video
    • Release statusEnterprise beta

      Seedance release tracker ↗Checked 2026-08-05

    Public comparative benchmark scores were unavailable at the check date.

  • MiniMax H3

    MiniMaxMiniMax

    open weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 15s video
    in · textin · imagein · audioin · videoout · video
    • Release capability2K · up to 15s

      MiniMax: H3 ↗Checked 2026-08-05

    Open Hailuo-line model with native stereo audio generation.

  • FLUX 3 Image

    Black Forest Labs

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Early access
    in · textin · imagein · audioin · videoout · image
    • Preliminary image evaluationImproved vs FLUX.2

      Black Forest Labs: FLUX 3 ↗Checked 2026-08-05

    Early image synthesis and editing surface from the FLUX 3 foundation model.

  • Veo 3.1

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 8s video
    in · textin · imageout · video
    • MovieGenBench T2V overall preferenceBest overall

      Google DeepMind: Veo 3.1 ↗Checked 2026-08-05

    Hosted video generation with native audio and reference-image control.

  • Sora 2 Pro

    OpenAIOpenAI

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 25s video
    in · textin · imageout · video
    • Artificial Analysis Video Arena1183 Elo

      Artificial Analysis: Sora 2 Pro ↗Checked 2026-08-05

    Higher-fidelity Sora tier for longer generated clips.

  • Kling 3.0

    KuaishouKuaishou

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 15s video
    in · textin · imagein · videoout · video
    • Artificial Analysis Video Arena (720p Standard)1211 Elo

      Artificial Analysis: Kling 3.0 Standard ↗Checked 2026-08-05

    Kling video generation family with native audio and reference controls.

  • Runway Gen-4.5

    Runway

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    Up to 10s video
    in · textin · imagein · videoout · video
    • Artificial Analysis T2V at launch1247 Elo

      Runway: Gen-4.5 ↗Checked 2026-08-05

    Runway video model with prompt and reference-based generation controls.

  • Wan 2.7

    AlibabaAlibaba

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    API release
    in · textin · imagein · videoout · video
    • Artificial Analysis T2V with audio1164 Elo

      Artificial Analysis: video leaderboard ↗Checked 2026-08-05

    Current hosted Wan video generation tier with audio output.

  • Wan 2.2 A14B

    AlibabaAlibaba

    open weights10 to 40B
    Params
    27B total · 14B active
    Context
    Open checkpoint
    in · textin · imageout · video
    • Release architecture14B active

      Alibaba: Wan 2.2 ↗Checked 2026-08-05

    Open video MoE with separate denoising experts.

  • Gemini Omni Flash

    GoogleGoogle

    closed weightsUndisclosed
    Params
    Undisclosed
    Context
    API release
    in · textin · imageout · video
    • Artificial Analysis T2V with audio1244 Elo

      Artificial Analysis: video leaderboard ↗Checked 2026-08-05

    Google video generation tier with synchronized audio output.