24 GB value workstation
Practical local entry point24 GB VRAM. A pragmatic used-GPU build for quantized local models without workstation pricing. $1,500.00 whole-system planning allowance.
Planning assumption ↗The short catalog
The catalog stays small and opinionated. Each entry links its evidence, so when a recommendation looks wrong to you, the receipts are one click away.
Open the planner6 compared
Planning envelopes for the machines the engine can recommend.
24 GB VRAM. A pragmatic used-GPU build for quantized local models without workstation pricing. $1,500.00 whole-system planning allowance.
Planning assumption ↗24 GB unified · 273 GB/s. A quiet, turnkey complete system at list price for MLX or llama.cpp local inference. $1,399.00 first-party list price.
Apple Mac mini newsroom ↗36 GB unified · 410 GB/s. A quiet, turnkey complete system at list price with more unified memory for MLX or llama.cpp serving. $1,999.00 first-party list price.
Apple Mac Studio newsroom ↗32 GB VRAM · 1,792 GB/s. Fast local inference for everyday models, backed by APIs only when the job merits it. $3,200.00 whole-system planning allowance.
NVIDIA RTX 5090 specifications ↗96 GB unified · 819 GB/s. A quiet, turnkey complete system at list price for capacity-first MLX or llama.cpp serving. $3,999.00 first-party list price.
Apple Mac Studio newsroom ↗128 GB unified · 273 GB/s. Prioritizes model capacity and privacy over the raw token speed of a consumer GPU. $4,699.00 first-party list price.
NVIDIA DGX Spark hardware ↗13 compared
Elastic routes for bursty workloads and the jobs that exceed local capability.
$0.14 input / $0.28 output per million tokens; uncached pricing.
DeepSeek API pricing ↗$0.20 input / $1.25 output per million tokens.
OpenAI API pricing ↗$0.30 input / $2.50 output per million tokens.
Gemini Developer API pricing ↗$1.00 input / $5.00 output per million tokens.
Claude API pricing ↗$0.20 input / $1.20 output per million tokens.
OpenAI API pricing ↗$0.44 input / $0.87 output per million tokens; uncached pricing.
DeepSeek API pricing ↗$0.75 input / $4.50 output per million tokens.
OpenAI API pricing ↗$1.50 input / $7.50 output per million tokens.
Gemini Developer API pricing ↗$2.00 input / $10.00 output per million tokens; introductory rate through August 31, 2026.
Claude API pricing ↗$2.00 input / $12.00 output per million tokens.
OpenAI API pricing ↗$2.00 input / $12.00 output per million tokens; paid tier for prompts up to 200K tokens.
Gemini Developer API pricing ↗$5.00 input / $30.00 output per million tokens; never used as the bulk route.
OpenAI API pricing ↗$5.00 input / $25.00 output per million tokens; never used as the bulk route.
Claude API pricing ↗8 compared
Private utility and reasoning routes sized to hardware a single developer could own.
A small multimodal utility model for light local work on hardware you already have.
Google Gemma 4 overview ↗A compact local reasoning model for coding and document work.
OpenAI gpt-oss-20b model card ↗A compact local multimodal model for coding and document work.
Mistral Small 3.3 model card ↗The local multimodal generalist for 24–32 GB consumer GPUs.
Google Gemma 4 getting started ↗A dense local model for coding and research on 24 GB class hardware.
Qwen3.6 27B model card ↗A large local multimodal model for private research and creative work.
Meta Llama 4 Scout model card ↗A large local reasoning model for technical work that fits on an 80 GB GPU.
OpenAI gpt-oss-120b model card ↗A large local coding and reasoning option; viable only with aggressive quantization and careful serving.
Official DeepSeek model weights ↗Quick starts