Gemma 4 31B
24 GB footprint
~23 tok/s · bandwidth estimate
Hardware
Memory decides which models fit. Bandwidth decides how fast they answer. Street price decides whether it ever pays for itself. Here is every machine in the planner, measured on all three.
The three numbers
A model runs only if its weights fit in GPU memory or unified memory, with room left for context. A 4-bit model needs a little over half a gigabyte per billion parameters.
Each generated token reads the active weights once. Divide memory bandwidth by the active weight size and you have a close ceiling on tokens per second. Mixture-of-experts models read only their active experts, which is why they feel fast.
Cards sell far above list right now. We price the machine you can actually buy today, then the planner compares it with months of API bills.
Fit and speed
Estimated generation speed at batch size one, assuming 60% of peak memory bandwidth, which is typical for a well-tuned llama.cpp, MLX, or TensorRT setup. A dash means the model does not fit.
| Machine | Gemma 4 E4B4 GB | gpt-oss-20b16 GB | Mistral Small 3.216 GB | Gemma 4 26B A4B18 GB | Qwen3.8 27B20 GB | Gemma 4 31B24 GB | Llama 4 Scout65 GB | gpt-oss-120b80 GB | DeepSeek V4 Flash 073190 GB |
|---|---|---|---|---|---|---|---|---|---|
| Used RTX 3090 class24 GB · 936 GB/s · ~$2,200 | 140 | 187 | 35 | 187 | 28 | 23 | — | — | — |
| M5 Pro Mac mini24 GB · 307 GB/s · ~$1,699 | 46 | 61 | 12 | 61 | 9 | 8 | — | — | — |
| RTX 5090 class32 GB · 1792 GB/s · ~$8,000 | 269 | 358 | 67 | 358 | 54 | 45 | — | — | — |
| M5 Max Mac Studio36 GB · 460 GB/s · ~$2,499 | 69 | 92 | 17 | 92 | 14 | 12 | — | — | — |
| M5 Ultra Mac Studio96 GB · 1200 GB/s · ~$5,499 | 180 | 240 | 45 | 240 | 36 | 30 | 72 | 180 | 144 |
| RTX PRO 6000 class96 GB · 1792 GB/s · ~$20,000 | 269 | 358 | 67 | 358 | 54 | 45 | 108 | 269 | 215 |
| DGX Spark128 GB · 273 GB/s · ~$5,000 | 41 | 55 | 10 | 55 | 8 | 7 | 16 | 41 | 33 |
| DGX Spark + RTX 5090128 GB · 1792 GB/s · ~$13,000 | 269 | 358 | 67 | 358 | 54 | 45 | 108 | 269 | 215 |
The machines
24GB
A used RTX 3090 (about $1,465 on the U.S. used market in September 2026) in a basic host, for quantized local models without workstation pricing.
24GB
A quiet, turnkey complete system at list price for MLX or llama.cpp local inference.
32GB
Fast local inference for everyday models. The card alone sells for about $6,800 new online, more than three times its $1,999 list price.
36GB
A quiet, turnkey complete system at list price with more unified memory for MLX or llama.cpp serving.
96GB
A quiet, turnkey complete system at list price for capacity-first MLX or llama.cpp serving.
96GB
A whole-system allowance around the 96 GB workstation card, which now sells for about $18,000 new, pairing big-model capacity with consumer-flagship speed.
128GB
Prioritizes model capacity and privacy over raw token speed. The $4,699 Founders Edition is out of stock, so this uses the cheapest available unit.
128GB
Route large models to Spark and latency-sensitive work to the desktop GPU.
Try your own
Choose a machine from the list or enter your own memory and bandwidth. Text and video models are matched against what fits, with speeds wherever the bandwidth is known.
24 GB footprint
~23 tok/s · bandwidth estimate
20 GB footprint
~28 tok/s · bandwidth estimate
18 GB footprint · 3 GB active
~187 tok/s · bandwidth estimate
16 GB footprint · 3 GB active
~187 tok/s · bandwidth estimate
16 GB footprint
~35 tok/s · bandwidth estimate
4 GB footprint
~140 tok/s · bandwidth estimate
20 GB · Expect about two minutes of GPU time for each second of finished video on a consumer GPU.
12 GB · Expect about one minute of GPU time for each second of finished video with the fast distilled route.
Too big for 24 GB: Llama 4 Scout (65 GB), gpt-oss-120b (80 GB), DeepSeek V4 Flash 0731 (90 GB).
The full recommendation weighs price and privacy too. Open the planner for a stack that fits your constraints.