Skip to contentT/TTokens or Towers

Compare

LeaderboardEvery tracked model, ranked and priced side by sideBenchmarksWhat each test measures and who leads itPricesAPI, media, and hardware rates with a cost calculator

Plan

Stack plannerSix questions to a priced local and hosted stackShortlistOur picks for subscriptions, APIs, and hardwareEnterpriseSeat plans against self-hosting for a whole team

Build

HardwareGPUs and workstations, and which models they fitHostsWhere to rent open-weight inference, and for how muchHarnessesCoding agents and the models they drive

Lab

FindingsResults from our own evaluationsScanAudit a repo’s AI usage in the browserMethodWhere the numbers come from and how we check them
Data checked September 25, 2026
Data checked September 25, 2026
Plan a stack
T/TTokens or Towers

An independent guide to AI models, what they cost, and the hardware to run them. No account, no API key, no sponsors.

Compare

  • Leaderboard
  • Benchmarks
  • Prices

Plan

  • Stack planner
  • Shortlist
  • Enterprise

Build

  • Hardware
  • Hosts
  • Harnesses

Lab

  • Findings
  • Scan
  • Method
Data checked September 25, 2026 · Every number links to its sourceRSS · Source on GitHub

Hardware

What your own box can actually run.

Memory decides which models fit. Bandwidth decides how fast they answer. Street price decides whether it ever pays for itself. Here is every machine in the planner, measured on all three.

Check a machine →Plan a stack
Machines compared
8
Open models sized
9
Cheapest memory
$39per GB
Most memory
128GB

The three numbers

Ignore the spec sheet. Read these three.

  1. Memory is the ceiling

    A model runs only if its weights fit in GPU memory or unified memory, with room left for context. A 4-bit model needs a little over half a gigabyte per billion parameters.

  2. Bandwidth is the speed

    Each generated token reads the active weights once. Divide memory bandwidth by the active weight size and you have a close ceiling on tokens per second. Mixture-of-experts models read only their active experts, which is why they feel fast.

  3. Street price is the verdict

    Cards sell far above list right now. We price the machine you can actually buy today, then the planner compares it with months of API bills.

Fit and speed

Every machine, every open model, one grid.

Estimated generation speed at batch size one, assuming 60% of peak memory bandwidth, which is typical for a well-tuned llama.cpp, MLX, or TensorRT setup. A dash means the model does not fit.

  • Under 8 tok/s batch jobs only
  • 8–25 tok/s reading speed
  • 25–70 tok/s comfortable chat and agents
  • 70+ tok/s faster than you can read
Estimated tokens per second for each open model on each machine
MachineGemma 4 E4B4 GBgpt-oss-20b16 GBMistral Small 3.216 GBGemma 4 26B A4B18 GBQwen3.8 27B20 GBGemma 4 31B24 GBLlama 4 Scout65 GBgpt-oss-120b80 GBDeepSeek V4 Flash 073190 GB
Used RTX 3090 class24 GB · 936 GB/s · ~$2,200140187351872823———
M5 Pro Mac mini24 GB · 307 GB/s · ~$1,6994661126198———
RTX 5090 class32 GB · 1792 GB/s · ~$8,000269358673585445———
M5 Max Mac Studio36 GB · 460 GB/s · ~$2,499699217921412———
M5 Ultra Mac Studio96 GB · 1200 GB/s · ~$5,49918024045240363072180144
RTX PRO 6000 class96 GB · 1792 GB/s · ~$20,000269358673585445108269215
DGX Spark128 GB · 273 GB/s · ~$5,0004155105587164133
DGX Spark + RTX 5090128 GB · 1792 GB/s · ~$13,000269358673585445108269215

The machines

From a used gaming card to a workstation.

All hardware prices →
  • 24GB

    24 GB value workstation

    A used RTX 3090 (about $1,465 on the U.S. used market in September 2026) in a basic host, for quantized local models without workstation pricing.

    Price
    ~$2,200
    Bandwidth
    936 GB/s
    Per GB
    $92
    ResalePrices used RTX 3090 market ↗
  • 24GB

    Mac mini (M5 Pro, 24 GB)

    A quiet, turnkey complete system at list price for MLX or llama.cpp local inference.

    Price
    ~$1,699
    Bandwidth
    307 GB/s
    Per GB
    $71
    Apple Mac mini newsroom ↗
  • 32GB

    32 GB performance workstation

    Fast local inference for everyday models. The card alone sells for about $6,800 new online, more than three times its $1,999 list price.

    Price
    ~$8,000
    Bandwidth
    1792 GB/s
    Per GB
    $250
    videocardprices.com RTX 5090 tracker ↗
  • 36GB

    Mac Studio (M5 Max, 36 GB)

    A quiet, turnkey complete system at list price with more unified memory for MLX or llama.cpp serving.

    Price
    ~$2,499
    Bandwidth
    460 GB/s
    Per GB
    $69
    Apple Mac Studio newsroom ↗
  • 96GB

    Mac Studio (M5 Ultra, 96 GB)

    A quiet, turnkey complete system at list price for capacity-first MLX or llama.cpp serving.

    Price
    ~$5,499
    Bandwidth
    1200 GB/s
    Per GB
    $57
    Apple Mac Studio newsroom ↗
  • 96GB

    96 GB Blackwell workstation

    A whole-system allowance around the 96 GB workstation card, which now sells for about $18,000 new, pairing big-model capacity with consumer-flagship speed.

    Price
    ~$20,000
    Bandwidth
    1792 GB/s
    Per GB
    $208
    videocardprices.com RTX PRO 6000 Blackwell tracker ↗
  • 128GB

    DGX Spark capacity appliance

    Prioritizes model capacity and privacy over raw token speed. The $4,699 Founders Edition is out of stock, so this uses the cheapest available unit.

    Price
    ~$5,000
    Bandwidth
    273 GB/s
    Per GB
    $39
    VideoCardz DGX Spark pricing report ↗
  • 128GB

    Capacity + speed local lab

    Route large models to Spark and latency-sensitive work to the desktop GPU.

    Price
    ~$13,000
    Bandwidth
    1792 GB/s
    Per GB
    $102
    NVIDIA DGX Spark hardware ↗

Try your own

Pick a machine, see what runs on it.

Choose a machine from the list or enter your own memory and bandwidth. Text and video models are matched against what fits, with speeds wherever the bandwidth is known.

Hardware

Text models

  • Gemma 4 31B

    24 GB footprint

    ~23 tok/s · bandwidth estimate

  • Qwen3.8 27B

    20 GB footprint

    ~28 tok/s · bandwidth estimate

  • Gemma 4 26B A4B

    18 GB footprint · 3 GB active

    ~187 tok/s · bandwidth estimate

  • gpt-oss-20b

    16 GB footprint · 3 GB active

    ~187 tok/s · bandwidth estimate

  • Mistral Small 3.2

    16 GB footprint

    ~35 tok/s · bandwidth estimate

  • Gemma 4 E4B

    4 GB footprint

    ~140 tok/s · bandwidth estimate

Video models

  • Wan 2.2 T2V A14B

    20 GB · Expect about two minutes of GPU time for each second of finished video on a consumer GPU.

  • LTX-Video

    12 GB · Expect about one minute of GPU time for each second of finished video with the fast distilled route.

Too big for 24 GB: Llama 4 Scout (65 GB), gpt-oss-120b (80 GB), DeepSeek V4 Flash 0731 (90 GB).

The full recommendation weighs price and privacy too. Open the planner for a stack that fits your constraints.