The short catalog

Every option should earn its slot.

The catalog stays small and opinionated. Each entry links its evidence, so when a recommendation looks wrong to you, the receipts are one click away.

Open the planner

6 compared

Hardware classes

Planning envelopes for the machines the engine can recommend.

24 GB value workstation

Practical local entry point

24 GB VRAM. A pragmatic used-GPU build for quantized local models without workstation pricing. $1,500.00 whole-system planning allowance.

Planning assumption

Mac mini (M4 Pro, 24 GB)

Turnkey Apple Silicon entry point

24 GB unified · 273 GB/s. A quiet, turnkey complete system at list price for MLX or llama.cpp local inference. $1,399.00 first-party list price.

Apple Mac mini newsroom

Mac Studio (M4 Max, 36 GB)

Unified-memory local workstation

36 GB unified · 410 GB/s. A quiet, turnkey complete system at list price with more unified memory for MLX or llama.cpp serving. $1,999.00 first-party list price.

Apple Mac Studio newsroom

32 GB performance workstation

Fast local default

32 GB VRAM · 1,792 GB/s. Fast local inference for everyday models, backed by APIs only when the job merits it. $3,200.00 whole-system planning allowance.

NVIDIA RTX 5090 specifications

Mac Studio (M3 Ultra, 96 GB)

High-capacity Apple Silicon workstation

96 GB unified · 819 GB/s. A quiet, turnkey complete system at list price for capacity-first MLX or llama.cpp serving. $3,999.00 first-party list price.

Apple Mac Studio newsroom

DGX Spark capacity appliance

Local model capacity

128 GB unified · 273 GB/s. Prioritizes model capacity and privacy over the raw token speed of a consumer GPU. $4,699.00 first-party list price.

NVIDIA DGX Spark hardware

13 compared

Hosted inference

Elastic routes for bursty workloads and the jobs that exceed local capability.

DeepSeek V4 Flash 0731

Value coding and batch route

$0.14 input / $0.28 output per million tokens; uncached pricing.

DeepSeek API pricing

Claude Haiku 4.5

Focused Claude economy route

$1.00 input / $5.00 output per million tokens.

Claude API pricing

DeepSeek V4 Pro

Stronger DeepSeek coding route

$0.44 input / $0.87 output per million tokens; uncached pricing.

DeepSeek API pricing

Claude Sonnet 5

Premium Claude generalist

$2.00 input / $10.00 output per million tokens; introductory rate through August 31, 2026.

Claude API pricing

Gemini 3.1 Pro

Premium Gemini multimodal route

$2.00 input / $12.00 output per million tokens; paid tier for prompts up to 200K tokens.

Gemini Developer API pricing

GPT-5.6 Sol

Difficult-task escalation

$5.00 input / $30.00 output per million tokens; never used as the bulk route.

OpenAI API pricing

Claude Opus 5

Premium difficult-task escalation

$5.00 input / $25.00 output per million tokens; never used as the bulk route.

Claude API pricing

8 compared

Local models

Private utility and reasoning routes sized to hardware a single developer could own.

Gemma 4 E4B

Existing-machine utility model

A small multimodal utility model for light local work on hardware you already have.

Google Gemma 4 overview

Qwen3.6 27B

24 GB dense coding generalist

A dense local model for coding and research on 24 GB class hardware.

Qwen3.6 27B model card

DeepSeek V4 Flash 0731 (experimental Q2)

Large local coding route

A large local coding and reasoning option; viable only with aggressive quantization and careful serving.

Official DeepSeek model weights

Quick starts

Inspect a complete preset.