Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub
Leaderboard/Sakana AI
Sakana AI

Sakana AI · Fugu

Fugu

Closed weightsSupersededproprietary

Original June 2026 Fugu orchestration tier; no longer listed on Sakana's pricing page after the Fugu Max launch.

AA Intelligence
—
Price
—
Output speed
—
Context window
—

Coding

  • SWE-bench Pro#15 of 31
    59%vendormini-swe-agentSakana AI: Fugu v1 benchmark table ↗

    Leader: Claude Opus 5.5 at 89.9%

  • Terminal-Bench 2.x#23 of 103
    80.2%vendorTerminal-Bench 2.1Sakana AI: Fugu v1 benchmark table ↗

    Leader: Claude Fable 5.1 at 91.4%

  • LiveCodeBench#1 of 56
    92.9%vendorSakana AI: Fugu v1 benchmark table ↗

    Leader: Fugu at 92.9%

  • SciCode#5 of 94
    60.1%vendorSakana AI: Fugu v1 benchmark table ↗

    Leader: Claude Opus 5.5 at 66.9%

Reasoning and knowledge

  • Humanity’s Last Exam#14 of 147
    47.2%vendorSakana AI: Fugu v1 benchmark table ↗

    Leader: Claude Opus 5.5 at 61.4%

  • GPQA Diamond#4 of 142
    95.5%vendorSakana AI: Fugu v1 benchmark table ↗

    Leader: GPT-6 Astra at 96.1%

Vision and long context

  • AA Long Context Reasoning#60 of 139
    74.7%vendor"Long Context Reasoning"Sakana AI: Fugu v1 benchmark table ↗

    Leader: Kimi K3 at 88.7%

Other published results

ResultScoreReported bySource
LiveCodeBench Pro87.8%vendorSakana AI: Fugu v1 benchmark table ↗
CharXiv Reasoning85.1%vendorSakana AI: Fugu v1 benchmark table ↗
tau3-Bench Banking21.7%vendorSakana AI: Fugu v1 benchmark table ↗
MRCRv286.6%vendorSakana AI: Fugu v1 benchmark table ↗

Specs

Released
June 22, 2026
Parameters
Undisclosed
Size class
Undisclosed
Input
text, image
Output
text
Context
Not published
Max output
Not published

    Will it fit your budget?

    The planner prices this model against local hardware and other APIs for your workload.

    Plan a stack →

    Compare with

    • Sakana AIFugu Ultra v2
    • Sakana AIFugu Max
    • OpenAIGPT-6 Astra52.7
    • OpenAIGPT-6 Sol47.5