Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/iFlytek
iFlytek

iFlytek · Spark X2.5

Spark-X2.5-4B

Open weightsGenerally availableNewApache-2.0

Small open-weight (Apache-2.0) iFlytek on-device model with a native 1M-token context, aimed at local agent and coding use; a 1.7B sibling was released alongside.

AA Intelligence
—
Price
Self-host
Output speed
—
Context window
1.05M tokens

Coding

  • SWE-bench Pro#27 of 31
    44.4%vendorthinking modeHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Claude Opus 5.5 at 89.9%

  • SWE-bench Verified#15 of 16
    41.6%vendorthinking modeHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Claude Sonnet 5 at 85.2%

  • SciCode#82 of 94
    34.7%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Claude Opus 5.5 at 66.9%

Agents and tool use

  • τ²-Bench#35 of 88
    75.1%vendordomain not statedHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: GLM-5.2 at 99.1%

  • BrowseComp#12 of 14
    40.9%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: GPT-6 Astra at 91.5%

  • MCP Atlas#9 of 9
    54.6%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Kimi K3 at 84.2%

Reasoning and knowledge

  • AIME 2026#7 of 13
    90.7%vendorthinking modeHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: ERNIE 5.1 at 99.6%

  • HMMT February 2026#6 of 8
    81.2%vendorHMMT Feb 2026, thinking modeHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: GLM-5.2 at 92.5%

  • GPQA Diamond#102 of 142
    67.4%vendor"GPQA" as labelled on card; thinking modeHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: GPT-6 Astra at 96.1%

  • Humanity’s Last Exam#86 of 147
    12.3%vendorthinking mode; tools not statedHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Claude Opus 5.5 at 61.4%

Vision and long context

  • AA Long Context Reasoning#87 of 139
    56.3%vendorvendor-run AA-LCRHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

    Leader: Kimi K3 at 88.7%

Other published results

ResultScoreReported bySource
BFCL-V465.1%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
tau3-bench30.4%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
MCP-Mark14.2%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
Workspace Bench31.2%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
VitaBench 2.025.2%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
SWE-bench Multilingual53.3%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
IMO-AnswerBench74.2%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
IFEval93%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
IFBench75%vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗
Gaokao 2026 (5 papers)133.4 points out of 150vendorHugging Face: XHToken/Spark-X2.5-4B (model card) ↗

Specs

Released
September 1, 2026
Parameters
4B
Size class
Under 10B
Input
text
Output
text
Context
1.05M tokens
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • iFlytekiFlytek Spark X2.5
    • OpenAIGPT-6 Astra52.7
    • OpenAIGPT-6 Sol47.5
    • OpenAIGPT-6 Luna37.3