Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/StepFun
StepFun

StepFun · Step 3

Step3 VL 10B

Open weightsGenerally availableApache-2.0

Small open-weight StepFun vision-language reasoning model for local/edge multimodal use.

AA Intelligence
7.7#115 of 142
Price
Self-host
Output speed
—
Context window
66K tokens

Composite indexes

  • Artificial Analysis Intelligence Index#115 of 142
    8independentv4.3.2 (AA estimate; independent evaluation forthcoming)Artificial Analysis: Step3 VL 10B ↗

    Leader: Claude Opus 5.5 at 58

Agents and tool use

  • τ²-Bench#70 of 88
    16.1%independentAA, telecomArtificial Analysis: Step3 VL 10B ↗

    Leader: GLM-5.2 at 99.1%

Reasoning and knowledge

  • Humanity’s Last Exam#100 of 147
    10.8%independentAA, no toolsArtificial Analysis: Step3 VL 10B ↗

    Leader: Claude Opus 5.5 at 61.4%

  • GPQA Diamond#99 of 142
    69%independentAAArtificial Analysis: Step3 VL 10B ↗

    Leader: GPT-6 Astra at 96.1%

  • CritPt#98 of 139
    0%independentAAArtificial Analysis: Step3 VL 10B ↗

    Leader: GPT-5.6 Sol at 32.3%

Vision and long context

  • AA Long Context Reasoning#128 of 139
    0%independentAAArtificial Analysis: Step3 VL 10B ↗

    Leader: Kimi K3 at 88.7%

  • MMMU-Pro#53 of 72
    64%independentAAArtificial Analysis: Step3 VL 10B ↗

    Leader: Claude Opus 5.5 at 87.7%

Other published results

ResultScoreReported bySource
Terminal-Bench Hard (AA) · AA5.3%independentArtificial Analysis: Step3 VL 10B ↗
IFBench (AA) · AA50.2%independentArtificial Analysis: Step3 VL 10B ↗
AA-Omniscience Index · AA-59 index (-100..100)independentArtificial Analysis: Step3 VL 10B ↗

Specs

Released
January 20, 2026
Parameters
10.2B
Size class
10 to 40B
Input
text, image
Output
text
Context
66K tokens
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • PerplexitySonar (Sonar API)7.7
    • GoogleGemma 4 E2B7.8
    • TIIFalcon-H1R-7B7.8
    • PerplexitySonar Pro (Sonar API)7.6