Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/MBZUAI
MBZUAI

MBZUAI · K2 Horizon

K2 Horizon 375B A23B

Open weightsGenerally availableNewApache-2.0

Flagship of MBZUAI IFM's fully open K2 Horizon family (Sept 2026): 375B/23B-active MoE with 512K context, released with weights, data and training code.

AA Intelligence
30.5#38 of 142
Price
Self-host
Output speed
—
Context window
524K tokens

Composite indexes

  • Artificial Analysis Intelligence Index#38 of 142
    31independentIntelligence Index v4.3; reasoningArtificial Analysis: K2 Horizon 375B A23B ↗

    Leader: Claude Opus 5.5 at 58

Coding

  • SciCode#62 of 94
    42.9%independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗
    42.7%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: Claude Opus 5.5 at 66.9%

  • Terminal-Bench 4.0#42 of 88
    1.5%independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗

    Leader: Claude Mythos 5.1 at 60.9%

  • Terminal-Bench 2.x#33 of 103
    71.9%independentTerminal-Bench 2.1; reasoningArtificial Analysis: K2 Horizon 375B A23B ↗
    70.2%vendorTerminal-Bench 2.1Hugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: Claude Fable 5.1 at 91.4%

  • SWE-bench Pro#28 of 31
    42.6%vendorstrict (no internet)Hugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: Claude Opus 5.5 at 89.9%

Agents and tool use

  • GDPval-AA#34 of 103
    1441vendorvendor-runHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
    1349independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗

    Leader: Claude Opus 5.5 at 1846

  • BrowseComp#9 of 14
    72.8%vendorDiscard-all@95k context managementHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: GPT-6 Astra at 91.5%

Reasoning and knowledge

  • GPQA Diamond#49 of 142
    87.3%independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗
    87.3%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: GPT-6 Astra at 96.1%

  • Humanity’s Last Exam#54 of 147
    32%independentno tools; reasoningArtificial Analysis: K2 Horizon 375B A23B ↗
    32%vendorno toolsHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: Claude Opus 5.5 at 61.4%

  • CritPt#51 of 139
    8.6%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
    4.6%independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗

    Leader: GPT-5.6 Sol at 32.3%

Vision and long context

  • AA Long Context Reasoning#32 of 139
    80%independentreasoningArtificial Analysis: K2 Horizon 375B A23B ↗
    76%vendorvendor-runHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

    Leader: Kimi K3 at 88.7%

Other published results

ResultScoreReported bySource
tau2-Bench Banking (AA) · reasoning34.2%independentArtificial Analysis: K2 Horizon 375B A23B ↗
AA-Omniscience Index · reasoning-3 indexindependentArtificial Analysis: K2 Horizon 375B A23B ↗
tau3-Bench Banking34%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
Toolathlon Verified65.3%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
AutomationBench · public set25.3%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
APEX-Agents · pass@1; text-only subset24.8%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
MCPMark67.7%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
WildClawBench · English text-only subset50.9%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
SWE-Atlas-QnA · strict (no internet)48.4%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
AA-Omniscience Accuracy · vendor-run23%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗
AA-Omniscience Non-Hallucination · vendor-run74.7%vendorHugging Face: IFM/K2-Horizon-375B-A23B model card ↗

Specs

Released
September 3, 2026
Parameters
375B total, 23B active
Size class
100 to 500B
Input
text
Output
text
Context
524K tokens
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • GoogleGemini 3.1 Pro Preview29.7
    • AlibabaQwen3.7-Max29
    • MiniMaxMiniMax-M329
    • OpenAIGPT-5.3 Codex32.5