Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/DeepSeek
DeepSeek

DeepSeek · DeepSeek V4

DeepSeek V4 Flash 0731

Open weightsSupersededNewMIT

Previous DeepSeek Flash tier (284B total / 13B active); superseded by V4.1 Flash on 2026-09-10, and its API name now routes to V4.1 Flash. Weights remain downloadable.

AA Intelligence
34#32 of 142
Price
Self-host
Output speed
—
Context window
1M tokens

Composite indexes

  • Artificial Analysis Intelligence Index#32 of 142
    34independentv4.3.2, reasoning, max effortArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Claude Opus 5.5 at 58

Coding

  • Terminal-Bench 4.0#28 of 88
    12.1%independentreasoning, max effort, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Claude Mythos 5.1 at 60.9%

  • Terminal-Bench 2.x#27 of 103
    82.7%vendorTerminal-Bench 2.1DeepSeek-V4-Pro-0813 model card (Hugging Face) ↗
    78.7%independentTerminal-Bench 2.1, reasoning, max effort, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Claude Fable 5.1 at 91.4%

  • SciCode#39 of 94
    50.3%independentreasoning, max effort, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Claude Opus 5.5 at 66.9%

Agents and tool use

  • GDPval-AA#27 of 103
    1427independentGDPval-AA v2.1, reasoning, max effortArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Claude Opus 5.5 at 1846

Reasoning and knowledge

  • GPQA Diamond#35 of 142
    90.8%independentreasoning, max effort, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: GPT-6 Astra at 96.1%

  • Humanity’s Last Exam#39 of 147
    51.5%vendorwith toolsDeepSeek-V4-Pro-0813 model card (Hugging Face) ↗
    38.6%independentreasoning, max effort, no tools, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗
    37.8%vendorno toolsDeepSeek-V4-Pro-0813 model card (Hugging Face) ↗

    Leader: Claude Opus 5.5 at 61.4%

  • CritPt#28 of 139
    16.6%independentreasoning, max effort, AA-runArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: GPT-5.6 Sol at 32.3%

Vision and long context

  • AA Long Context Reasoning#33 of 139
    79.7%independentreasoning, max effort, AA-LCRArtificial Analysis: DeepSeek V4 Flash 0731 (Reasoning, Max Effort) ↗

    Leader: Kimi K3 at 88.7%

Human preference

  • LMArena Text#41 of 54
    1436independentarena entry "deepseek-v4-flash", 48,887 votes, checked 2026-09-25Arena (LMArena) text leaderboard ↗

    Leader: Claude Fable 5 at 1506

  • LMArena WebDev#25 of 54
    1579independentarena entry "deepseek-v4-flash-high", 4,724 votes, checked 2026-09-25Arena (LMArena) WebDev/Code leaderboard ↗

    Leader: Claude Opus 5.5 at 1818

Other published results

ResultScoreReported bySource
DeepSWE54.4%vendorDeepSeek-V4-Pro-0813 model card (Hugging Face) ↗
Agents' Last Exam25.2%vendorDeepSeek-V4-Pro-0813 model card (Hugging Face) ↗

Specs

Released
July 31, 2026
Parameters
284B total, 13B active
Size class
100 to 500B
Input
text
Output
text
Context
1M tokens
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • GoogleGemini 3.6 Flash34
    • AlibabaQwen3.8-27B34
    • Z.aiGLM-5.234
    • MetaMuse Spark 1.133.7