T/TTokens or Towers
PlannerShortlistScanPricesBenchmarksHarnessesHostsEnterprise
SITEMAP.XML · checked August 4, 2026
Tokens or Towers prototypeEvery recommendation shows its work.

The shortlist

What we would tell a friend.

This is the short answer: a few picks in each category with one reason each. The reference pages hold the whole field when you need to dig.

Open the planner →

Subscriptions

For an individual, a flat monthly plan usually beats buying the same traffic at API rates. Twenty dollars of plan buys far more inference than twenty dollars of tokens.

Claude Max 5x

Anthropic

The hundred dollar Max tier is the working default for daily agentic coding on Claude.

$100.00 / month

Anthropic plan pricing ↗

Claude Max 20x

Anthropic

Twenty times the Pro allowance for the weeks when the agents never sleep.

$200.00 / month

Anthropic plan pricing ↗

ChatGPT Pro 5x

OpenAI

The mid OpenAI tier when Codex is your daily driver and Plus limits pinch.

$100.00 / month

ChatGPT plan pricing ↗

ChatGPT Pro 20xEditor's choice

OpenAI

Maximum Codex access that would cost thousands a month at API rates.

$200.00 / month

ChatGPT plan pricing ↗

Kimi Code Vivace

Moonshot AI

The biggest Kimi Code quota, with open-weight K3 behind it keeping the exit door open.

$159.00 / month

Kimi Code plans ↗

Z.ai Max

Z.ai

GLM-5.2 at fourteen times Lite usage for well under the frontier plan prices.

$117.60 / month

Z.ai subscription plans ↗

Harnesses

The tool you live in matters as much as the model. These are the four we reach for.

Claude Code

Anthropic

The reference terminal agent for a Claude plan.

Claude Pro/Max subscription or Anthropic API keys.

Product site ↗

Codex CLI

OpenAI

Open source, fast, and the natural pair for a ChatGPT plan.

ChatGPT subscription auth or OpenAI API usage.

Source repo ↗

VS Code

Microsoft

The default editor that almost everything else plugs into.

Editor is free; agent capability comes from extensions.

Source repo ↗

T3 CodeEditor's choice

ping.gg

One control surface over whichever provider agents you already run.

Free control surface; you pay for whichever provider CLIs you already run.

Source repo ↗

Hosted models

When you do pay per token, these six cover the field from cheap bulk work to the strongest thing money buys.

Claude Fable 5Editor's choice

escalation · hosted

The strongest generally available model right now. Save it for work that earns the rate.

$10.00 in / $50.00 out per 1M tokens

Claude API pricing ↗

GPT-5.6 Sol

escalation · hosted

The OpenAI flagship for long-horizon agentic work.

$5.00 in / $30.00 out per 1M tokens

OpenAI API pricing ↗

GPT-5.6 Luna

generalist · hosted

The cheap multimodal generalist the planner keeps picking on its own.

$0.20 in / $1.20 out per 1M tokens

OpenAI API pricing ↗

Kimi K3

generalist · hosted

Open weights at frontier scale, so the hosted route never locks you in.

$3.00 in / $15.00 out per 1M tokens

Kimi API pricing ↗

GLM-5.2

generalist · hosted

The current GLM flagship, priced under the western frontier tiers.

$1.40 in / $4.40 out per 1M tokens

Z.ai model pricing ↗

DeepSeek V4 Flash 0731

economy · hosted

The inexpensive default when agents and batch jobs burn tokens by the hundred million.

$0.14 in / $0.28 out per 1M tokens

DeepSeek API pricing ↗

Hosts

Hosts are where tokens leave the building. A speed specialist, a universal key, and the gateway we would build on.

Cerebras Inference

Cerebras

The speed play: open models served faster than anyone else runs them.

Positions on wafer-scale chips for very high tokens-per-second on supported models.

Pricing ↗

OpenRouter

OpenRouter

One key across nearly every backend that matters.

Pricing ↗

Vercel AI GatewayEditor's choice

Vercel

Pass-through pricing behind one gateway, living where the deploys already happen.

Pricing ↗

Hardware

Buy the smallest box that runs the local models you care about, and let capacity buy privacy at scale.

DGX Spark capacity applianceEditor's choice

DGX Spark

The capacity appliance the local half of this site is built around.

~$4,699.00

NVIDIA DGX Spark hardware ↗

96 GB Blackwell workstation

RTX PRO 6000 class

Ninety-six gigabytes at consumer-flagship bandwidth for a serious local lab.

~$9,500.00

NVIDIA RTX PRO 6000 Blackwell ↗

32 GB performance workstation

RTX 5090 class

The fast local default when your models fit in 32 GB.

~$3,200.00

NVIDIA RTX 5090 specifications ↗

Video

Video stays optional. One route for quality, one cheap enough to draft with all day.

Seedance 2.5

Hosted video

The quality pick per second among the hosted video routes.

Dreamina Seedance 2.5 ↗

MiniMax H3Editor's choice

Hosted video

Cheap enough per second to draft with, good enough to ship.

$0.08 per second of 1080p

MiniMax API pricing ↗

Quick starts

Inspect a complete preset.

The subscription developerFor a single developer, a couple of two hundred dollar plans buy the best frontier-token value on the market today. Nobody runs a Kimi K3 at home; the machines that hold a 2.8T model cost more than years of subscriptions.View stack →Private local labFor the developer who wants local on purpose. A capacity-first one-person setup for local-only coding, vision, and documents.View stack →Enterprise: plans or on-premPer-developer math for a team decision. The three routes are the argument: metered cloud, the flat-plan route, and an on-prem tower running open models. Multiply by your seat count, and note that enterprise seat pricing usually lands between the plan and API lines.View stack →

Pick a machine, see what runs on it.

Match local models to hardware memory before you buy the tower.

Hardware fit

Pick a machine. See what runs.

Local models that fit this memory, with speeds from bandwidth when we have a figure.

Hardware

Text models

  • Gemma 4 31B

    24 GB footprint

    ~23 tok/s · bandwidth estimate

  • Qwen3.6 27B

    20 GB footprint

    ~28 tok/s · bandwidth estimate

  • Gemma 4 26B A4B

    18 GB footprint · 3 GB active

    ~187 tok/s · bandwidth estimate

  • gpt-oss-20b

    16 GB footprint · 3 GB active

    ~187 tok/s · bandwidth estimate

  • Mistral Small 3.3

    16 GB footprint

    ~35 tok/s · bandwidth estimate

  • Gemma 4 E4B

    4 GB footprint

    ~140 tok/s · bandwidth estimate

Video models

  • Wan 2.2 T2V A14B

    20 GB · Expect about two minutes of GPU time for each second of finished video on a consumer GPU.

  • LTX-Video

    12 GB · Expect about one minute of GPU time for each second of finished video with the fast distilled route.

Too big for 24 GB: Llama 4 Scout (65 GB), gpt-oss-120b (80 GB), DeepSeek V4 Flash 0731 (90 GB).

The full recommendation weighs price and privacy too. Open the planner for a stack that fits your constraints.