Benchmarks reference
The evidence, filterable. A dated snapshot of provider claims and independent evaluations. Every claim links to its source. All measurements come from their publishers.
Open the planner → Size buckets follow total parameter count, from under 10B up to the trillion-scale ranges, and each entry also lists its active parameters where published. Closed models with unpublished counts appear as Undisclosed. Scores reflect claims checked through 2026-08-05. Eval setups vary, so compare scores only within the same suite.
Provider Alibaba Amazon Anthropic Baidu Black Forest Labs ByteDance Cohere DeepSeek Google Kuaishou Meta Microsoft MiniMax Mistral Moonshot AI NVIDIA OpenAI Runway Sakana AI Thinking Machines Lab Z.ai xAI
Weights Any open closed
Size class Any Under 10B 10 to 40B 100 to 500B 500B to 1T 1 to 2T 2 to 4T Undisclosed
Input types text image audio video
Output types text image audio video
Reset filters
Params Undisclosed
Context 1M tokens in · text in · image out · text
Fast/affordable tier of the GPT-5.6 family. Parameter count undisclosed.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Mid GPT-5.6 hosted tier between Luna and Sol. Parameter count undisclosed.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Flagship GPT-5.6 model for long-horizon agentic work. Parameter count undisclosed.
Params Undisclosed
Context 400K tokens in · text in · image out · text
Compact GPT-5.4 generalist for coding and document work. Parameter count undisclosed.
Params Undisclosed
Context 400K tokens in · text in · image out · text
Lowest-cost GPT-5.4 text tier. Parameter count undisclosed.
Params 284B total · 13B active
Context 1M tokens in · text out · text
MIT open weights; MoE with ~13B active. Official post-training release superseding the preview.
Params Undisclosed
Context 1M tokens in · text out · text
Open-weight DeepSeek Pro tier for harder coding and research. Parameter count varies by reporting source.
Params 4B
Context 128K tokens in · text in · image in · audio in · video out · text
Edge/mobile multimodal utility model from the Gemma 4 family.
Params 26B total · 3.8B active
Context 256K tokens in · text in · image in · video out · text
MoE local generalist: ~26B total / ~3.8B active. Fits 24–32 GB consumer GPUs when quantized.
Params 31B
Context 256K tokens in · text in · image in · video out · text
Dense Gemma 4 flagship for higher-quality local multimodal work.
Params Undisclosed
Context 1M tokens in · text in · image in · audio in · video out · text
Hosted Gemini Pro multimodal tier. Parameter count undisclosed.
Params Undisclosed
Context 1M tokens in · text in · image in · audio in · video out · text
Fast multimodal Flash generalist on the current Gemini 3.6 line.
Params Undisclosed
Context 1M tokens in · text in · image in · audio in · video out · text
Low-cost multimodal Flash-Lite route for high-volume document and vision work.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Premium Claude generalist. Parameter count undisclosed.
Params Undisclosed
Context 200K tokens in · text in · image out · text
Economy Claude route for focused coding and document work.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Anthropic frontier model; parameter count undisclosed. Some evals fall back to Opus 4.8 under safety guardrails.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Prior Anthropic flagship still used as a fallback and comparison baseline.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Current Claude Opus API tier. Parameter count undisclosed.
Params 21B total · 3.6B active
Context 128K tokens in · text out · text
Open-weight MoE (~21B total / 3.6B active). Entry-point coding model for 16 GB-class machines.
Params 117B total · 5.1B active
Context 128K tokens in · text out · text
Open-weight MoE (~117B total / 5.1B active). Capacity-class coding route for 80 GB+ machines.
Params 109B total · 17B active
Context 10M tokens in · text in · image out · text
Open-weight MoE (17B active / 109B total) with a 10M-token context window.
Params 400B total · 17B active
Context 1M tokens in · text in · image out · text
Larger Llama 4 MoE (17B active, 128 experts). Total parameter count commonly reported near 400B.
Params 27B
Context 262K tokens in · text in · image in · video out · text
Open-weight mid-size Qwen tier; strong value for local or self-hosted runs.
Params Undisclosed
Context 1M tokens in · text out · text
Hosted Qwen Max tier. Parameter count undisclosed.
Params Undisclosed
Context 1M tokens in · text in · image in · video out · text
Hosted multimodal Qwen Plus tier. Parameter count undisclosed.
Params 35B total · 3B active
Context 256K tokens in · text in · image in · video out · text
Hosted alias for the open Qwen3.6 35B-A3B checkpoint.
Params 35B total · 3B active
Context 256K tokens in · text in · image in · video out · text
Open MoE release behind the Qwen3.6 Flash API tier.
Params 2.4T total · 95B active
Context Not published in · text in · image out · text
The first Max-class Qwen with announced open weights, scheduled for release the week after the August 2 launch. 2.4T total with 95B active. This entry flips to open the day the weights land.
Params 235B total · 22B active
Context 262K tokens in · text out · text
Open-weight MoE flagship of the Qwen3.7 line.
Params 24B
Context 128K tokens in · text in · image out · text
Dense 24B multimodal open weights. Scores cite the published Small 3.2 instruct card for the same size class.
Params Undisclosed
Context 1M tokens in · text in · image in · audio in · video out · text
Hosted multimodal Flash tier. Index scores shifted between methodology versions; both dated claims are linked.
Params 2.8T
Context 1M tokens in · text out · text
Open-weight long-horizon coding/reasoning model, commonly reported near 2.8T total parameters.
Params 1T total · 32B active
Context 256K tokens in · text in · image in · video out · text
Open coding model with 1T total parameters and 32B active per token.
Params 744B total · 40B active
Context 200K tokens in · text out · text
Open-weight MoE (~744B total / 40B active). Useful open frontier baseline beside Kimi.
Params 744B total · 40B active
Context 200K tokens in · text out · text
Successor release on the GLM-5 base with stronger agentic coding results.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Prior Grok API tier retained for comparison. Parameter count undisclosed.
Params Undisclosed
Context Not published in · text out · text
Coding model available through Grok Build. Parameter count undisclosed.
Params Undisclosed
Context 500K tokens in · text in · image out · text
Current xAI flagship reasoning tier. Parameter count undisclosed.
Params 675B total · 41B active
Context 256K tokens in · text in · image out · text
Open-weight MoE (~675B total / 41B active). Current Mistral Large API tier.
Params Undisclosed
Context 300K tokens in · text in · image in · video out · text
Bedrock multimodal Pro tier with 300K context. Parameter count undisclosed.
Params 111B
Context 256K tokens in · text out · text
Cohere enterprise Command tier. Open weights under a non-commercial research license.
Params 230B total · 10B active
Context 200K tokens in · text out · text
Open-weight MoE (~230B total / 10B active) aimed at agentic coding.
Params Undisclosed
Context 256K tokens in · text in · image out · text
ByteDance Seed coding/reasoning API model. Parameter count undisclosed.
Params 300B total · 47B active
Context 128K tokens in · text out · text
Open-weight MoE (~300B total / 47B active) from Baidu Ernie.
Params 975B total · 41B active
Context 1M tokens in · text in · image in · audio out · text
Open MoE model with 975B total parameters and 41B active per token.
Params 276B total · 12B active
Context 256K tokens in · text in · image in · audio out · text
Preview checkpoint with 276B total parameters and 12B active. Full weights are pending.
Params 30B total · 3B active
Context 256K tokens in · text in · image in · audio in · video out · text
Open omni model with 30B total parameters and 3B active per token.
Params 120B total · 12B active
Context 1M tokens in · text out · text
Open agentic MoE with 120B total parameters and 12B active per token.
Params 550B total · 55B active
Context 1M tokens in · text out · text
Open Nemotron flagship with 550B total parameters and 55B active per token.
Params 119B total · 6B active
Context 256K tokens in · text in · image out · text
Apache 2.0 MoE with 119B total parameters and 6B active per token.
Params 14B
Context 16K tokens in · text out · text
Dense 14B open model focused on compact reasoning workloads.
Params 3.8B
Context 128K tokens in · text out · text
Dense 3.8B open model for local text inference.
Params 5.6B
Context 128K tokens in · text in · image in · audio out · text
Open 5.6B model for multimodal understanding.
Params 428B total · 23B active
Context 1M tokens in · text in · image in · video out · text
Open multimodal MoE with 428B total parameters and 23B active per token.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Orchestration model that routes hard tasks to a pool of expert models. Scores are Sakana-reported and independent verification is pending.
Params Undisclosed
Context 1M tokens in · text in · image out · text
Lower-latency orchestration tier that bills at the routed model rate. Scores are Sakana-reported and independent verification is pending.
Params Undisclosed
Context Up to 15s video in · text in · image in · audio in · video out · video
Native multimodal video generation with synchronized audio.
Params Undisclosed
Context Enterprise beta in · text in · image in · audio in · video out · video
Public comparative benchmark scores were unavailable at the check date.
Params Undisclosed
Context Up to 15s video in · text in · image in · audio in · video out · video
Open Hailuo-line model with native stereo audio generation.
Params Undisclosed
Context Early access in · text in · image in · audio in · video out · image
Early image synthesis and editing surface from the FLUX 3 foundation model.
Params Undisclosed
Context Up to 8s video in · text in · image out · video
Hosted video generation with native audio and reference-image control.
Params Undisclosed
Context Up to 25s video in · text in · image out · video
Higher-fidelity Sora tier for longer generated clips.
Params Undisclosed
Context Up to 15s video in · text in · image in · video out · video
Kling video generation family with native audio and reference controls.
Params Undisclosed
Context Up to 10s video in · text in · image in · video out · video
Runway video model with prompt and reference-based generation controls.
Params Undisclosed
Context API release in · text in · image in · video out · video
Current hosted Wan video generation tier with audio output.
Params 27B total · 14B active
Context Open checkpoint in · text in · image out · video
Open video MoE with separate denoising experts.
Params Undisclosed
Context API release in · text in · image out · video
Google video generation tier with synchronized audio output.