The sensible default
The best balance of capability, privacy, and money for the constraints you chose.
- Hardware
- ~$1,500
- API / month
- $3–$5
- Runs locally
- 80%
- Breaks even
- Tokens win
A pragmatic used-GPU build for quantized local models without workstation pricing.
Planning assumption ↗- 24 GB is a practical entry point for useful quantized local models.
- Routine work stays on the economical route while difficult jobs have a clear fallback.
Reality check
- Treat hardware prices as planning allowances and check a live listing first.
- API estimate assumes 70% input / 30% output and excludes caching discounts.
- Token speeds come from memory-bandwidth math at batch size 1. Measure before you rely on them.