The sensible default
The best balance of capability, privacy, and money for the constraints you chose.
- Hardware
- ~$3,200
- Cloud / month
- $15–$24
- Runs locally
- 80%
- Breaks even
- Tokens winThe hardware never recoups its cost at this volume. Tokens stay cheaper.
Fast local inference for everyday models, backed by APIs only when the job merits it.
NVIDIA RTX 5090 specifications ↗- Pairs high local token speed with enough VRAM for a capable multimodal model.
- Routine work stays on the economical route while difficult jobs have a clear fallback.
- Ranked on a measured intelligence index of 37.
Reality check
- Treat hardware prices as planning allowances and check a live listing first.
- API estimate assumes 70% input / 30% output and excludes caching discounts.
- Sustained heavy volume can outrun personal plan limits.
- Token speeds come from memory-bandwidth math at batch size 1. Measure before you rely on them.