Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/Alibaba
Alibaba

Alibaba · Qwen3.8

Qwen3.8-Flash-Next

Open weightsGenerally availableNewcustom (Qwen Community License 1.0)

Open preview of the architecture planned for Qwen4 (125B MoE with 6B active, plus 51B n-gram embedding and 4B MTP parameters, about 180B total), and the base of the hosted Qwen3.8-Flash.

AA Intelligence
40#21 of 142
Price
Self-host
Output speed
54tok/s
Context window
262K tokens

Composite indexes

  • Artificial Analysis Intelligence Index#21 of 142
    40independentv4.3.2, reasoningArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Opus 5.5 at 58

Coding

  • Terminal-Bench 4.0#19 of 88
    25.3%independentreasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Mythos 5.1 at 60.9%

  • Terminal-Bench 2.x#12 of 103
    86.1%independentTerminal-Bench 2.1, reasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Fable 5.1 at 91.4%

  • SciCode#37 of 94
    50.6%independentreasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Opus 5.5 at 66.9%

  • SWE-bench Pro#9 of 31
    62.5%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗

    Leader: Claude Opus 5.5 at 89.9%

  • LiveCodeBench#2 of 56
    91.9%vendorLiveCodeBench v6Qwen3.8-Flash-Next model card (Hugging Face) ↗

    Leader: Fugu at 92.9%

Agents and tool use

  • GDPval-AA#10 of 103
    1612independentGDPval-AA v2.1, reasoningArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Opus 5.5 at 1846

Reasoning and knowledge

  • GPQA Diamond#22 of 142
    92.3%independentreasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗
    91.7%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗

    Leader: GPT-6 Astra at 96.1%

  • Humanity’s Last Exam#42 of 147
    38%independentreasoning, no tools, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗
    35.9%vendorno toolsQwen3.8-Flash-Next model card (Hugging Face) ↗

    Leader: Claude Opus 5.5 at 61.4%

  • CritPt#36 of 139
    11.1%independentreasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: GPT-5.6 Sol at 32.3%

Vision and long context

  • AA Long Context Reasoning#34 of 139
    79.7%independentreasoning, AA-LCRArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Kimi K3 at 88.7%

  • MMMU-Pro#18 of 72
    79.8%independentreasoning, AA-runArtificial Analysis: Qwen3.8-Flash-Next ↗

    Leader: Claude Opus 5.5 at 87.7%

Human preference

  • LMArena WebDev#9 of 54
    1636independentarena entry "qwen3.8-flash-next (preliminary)", 4,429 votes, checked 2026-09-25Arena (LMArena) WebDev/Code leaderboard ↗

    Leader: Claude Opus 5.5 at 1818

Other published results

ResultScoreReported bySource
DeepSWE 1.158.7%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗
SWE-bench Multilingual81%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗
IFBench81.3%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗
Toolathlon Verified · Pass@173.5%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗
OSWorld 2.0 (binary) · partial-credit score 52.319.4%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗
AndroidWorld84.5%vendorQwen3.8-Flash-Next model card (Hugging Face) ↗

Specs

Released
August 26, 2026
Parameters
180B total, 6B active
Size class
100 to 500B
Input
text, image, video
Output
text
Context
262K tokens
Max output
Not published
First token
2.86s
  • Speed: Artificial Analysis: Qwen3.8-Flash-Next ↗

Will it fit your budget?

The planner prices this model against local hardware and other APIs for your workload.

Plan a stack →

Compare with

  • AlibabaQwen3.8-2.4T-A95B40
  • MetaMuse Spark 1.239.6
  • GoogleGemini 3.8 Flash40.9
  • GoogleGemini 3.7 Flash39.1