Skip to contentT/TTokens or Towers
01Planner02Shortlist03Scan04Prices05Benchmarks06Findings07Harnesses08Hosts09Enterprise
US · checked August 4, 2026
US · checked August 4, 2026
Tokens or Towers prototypeHow we workEvery recommendation shows its work.

Reference prices

Current rates.

Every row cites the provider or product page it came from, with the date we last checked it. Street estimates and planning allowances carry a leading ~ so you can tell them apart from list prices. Confirm against the live page before you spend anything.

Prices last checked August 5, 2026

Open the planner →

Tokens, translated

What a million tokens actually is

Providers price AI in tokens. These translations give the unit some ordinary scale.

One token
A chunk of text a model reads or writes. In English, one token is roughly three quarters of a word.
1,000 tokens
Roughly 750 words, about a page and a half.
1M tokens
Roughly 750,000 words, about seven paperback novels.
A typical chat
A question and answer runs a few hundred tokens.
Agents
One coding-agent task can read hundreds of thousands of tokens of code and logs, so a single busy session can pass 1M.
Light · 20M / month
Occasional agent use beside daily chat.
Regular · 100M / month
An agent working beside you most days.
Heavy · 500M / month
Agents running much of the day.

The calculator below turns a monthly token guess into an estimated bill.

Estimate a bill

Monthly cost estimator

Price your model usage.

This is an estimate. The token primer above explains what these volumes mean.

Model
Input
$0.14 / 1M
Output
$0.28 / 1M
Monthly usage per seat
Seats

Enter 1 to 500 seats. Values outside that range clamp when you leave the field.

Live estimate

DeepSeek V4 Flash 0731

Per seat monthly
$1.96
Total monthly at 1 seat
$1.96

10M in + 2M out, per seat

Reality check
  • These numbers are estimates based on published list rates.
  • Caching and batch discounts are excluded.
  • Enterprise contracts are negotiated and usually land below list.
  • The seat plan comparison uses public team rates for dedicated coding seats.

62 models

Hosted model APIs

Input and output rates are USD per million tokens unless a row note says otherwise.

Vercel AI Gateway passes through provider list prices with no markup. Alibaba's current Model Studio pricing and Token Plan allowlist stop at Qwen 3.7, so no Qwen 3.8 price is shown.

Hosted model API prices with source and check date
ProviderModelInput $/MOutput $/MContextSourceChecked
DeepSeekDeepSeek V4 Flash 0731Cache-miss (uncached) rates; cache hits are far cheaper.$0.14$0.281MDeepSeek API pricing ↗2026-08-05
DeepSeekDeepSeek V4 ProCache-miss rates from the DeepSeek pricing table.$0.44$0.871MDeepSeek API pricing ↗2026-08-05
OpenAIGPT-5.6 LunaShort-context standard rates; prompts over 272K are surcharged.$0.20$1.201MOpenAI API pricing ↗2026-08-05
OpenAIGPT-5.6 TerraShort-context standard rates from the OpenAI pricing table.$2.00$12.001MOpenAI API pricing ↗2026-08-05
OpenAIGPT-5.6 SolEscalation tier; short-context standard rates.$5.00$30.001MOpenAI API pricing ↗2026-08-05
OpenAIGPT-5.4 miniStandard rates from the OpenAI pricing table.$0.75$4.501MOpenAI API pricing ↗2026-08-05
OpenAIGPT-5.4 nanoLowest GPT-5.4 text tier on the OpenAI pricing page.$0.20$1.251MOpenAI API pricing ↗2026-08-05
AnthropicClaude Haiku 4.5Standard global Claude API rate.$1.00$5.00200KClaude API pricing ↗2026-08-05
AnthropicClaude Sonnet 5Current global Claude API rate.$2.00$10.001MClaude API pricing ↗2026-08-05
AnthropicClaude Opus 5Standard global Claude API rate.$5.00$25.001MClaude API pricing ↗2026-08-05
AnthropicClaude Fable 5Standard rates from the Claude pricing table; batch runs at half price.$10.00$50.001MClaude API pricing ↗2026-08-05
GoogleGemini 3.5 Flash-LitePaid-tier standard rates (text/image/video/audio input).$0.30$2.501MGemini Developer API pricing ↗2026-08-05
GoogleGemini 3.6 FlashPaid-tier standard rates; output includes thinking tokens.$1.50$7.501MGemini Developer API pricing ↗2026-08-05
GoogleGemini 3.1 ProPaid tier for prompts ≤200K tokens; longer prompts are $4/$18.$2.00$12.001MGemini Developer API pricing ↗2026-08-05
xAIGrok 4.3Standard rates for prompts up to 200K tokens.$1.25$2.501MxAI API pricing ↗2026-08-05
xAIGrok 4.5Standard rates for prompts up to 200K tokens.$2.00$6.00500KxAI API pricing ↗2026-08-05
xAIComposer 2.5Cursor standard rate. Fast mode costs $3 input and $15 output; xAI also offers the model in Grok Build.$0.50$2.50200KCursor Composer 2.5 pricing ↗2026-08-05
MistralMistral Large 3Standard La Plateforme API rates.$0.50$1.50256KMistral API pricing ↗2026-08-05
MistralMistral Medium 3.5Standard La Plateforme API rates.$1.50$7.50256KMistral API pricing ↗2026-08-05
MistralMistral Small 4Standard La Plateforme API rates.$0.15$0.60256KMistral API pricing ↗2026-08-05
Moonshot AIKimi K3Cache-miss rates from the official Kimi API table.$3.00$15.001MKimi API pricing ↗2026-08-05
Moonshot AIKimi K2.7 CodeCache-miss input rate. Cache hits cost $0.19 per million tokens.$0.95$4.00256KKimi K2.7 Code pricing ↗2026-08-05
Moonshot AIKimi K2.7 Code HighSpeedCache-miss input rate. Kimi lists about 180 output tokens per second and up to 260 in short contexts.$1.90$8.00256KKimi K2.7 Code pricing ↗2026-08-05
AlibabaQwen 3.7 MaxSingapore international list rates across the full context window.$2.50$7.501MAlibaba Cloud Model Studio pricing ↗2026-08-05
AlibabaQwen 3.7 Plus (up to 256K input)Singapore international list rates for prompts up to 256K tokens.$0.40$1.601MAlibaba Cloud Model Studio pricing ↗2026-08-05
AlibabaQwen 3.7 Plus (256K to 1M input)Singapore international list rates for requests above 256K input tokens.$1.20$4.801MAlibaba Cloud Model Studio pricing ↗2026-08-05
AlibabaQwen 3.6 FlashSingapore international rates for prompts up to 256K tokens.$0.25$1.501MAlibaba Cloud Model Studio pricing ↗2026-08-05
Z.aiGLM-5Standard cache-miss rates from the Z.ai pricing table.$1.00$3.20200KZ.ai model pricing ↗2026-08-05
Z.aiGLM-5.2Standard cache-miss rates from the Z.ai pricing table.$1.40$4.40200KZ.ai model pricing ↗2026-08-05
Z.aiGLM-4.5-AirStandard cache-miss rates from the Z.ai pricing table.$0.20$1.10128KZ.ai model pricing ↗2026-08-05
AmazonAmazon Nova ProBedrock standard on-demand rates in primary US regions.$0.80$3.20300KAmazon Bedrock pricing ↗2026-08-05
AmazonAmazon Nova LiteBedrock standard on-demand rates in primary US regions.$0.06$0.24300KAmazon Bedrock pricing ↗2026-08-05
CohereCommand AStandard production API rates.$2.50$10.00256KCohere Command A model card ↗2026-08-05
Thinking Machines LabInkling-SmallTinker serverless inference beta rate. Cached input costs $0.06 per million tokens.$0.30$1.20256KTinker models and pricing ↗2026-08-05
Thinking Machines LabInklingTinker serverless inference beta rate. Cached input costs $0.17 per million tokens.$1.00$4.05256KTinker models and pricing ↗2026-08-05
NVIDIANemotron 3 Nano 30B A3BBedrock standard on-demand rate in the primary US regions.$0.06$0.24256KAmazon Bedrock pricing ↗2026-08-05
NVIDIANemotron 3 Super 120B A12BBedrock standard on-demand rate in the primary US regions.$0.15$0.65256KAmazon Bedrock pricing ↗2026-08-05
NVIDIANemotron 3 Ultra 550B A55BTogether serverless list rate. Cached input costs $0.20 per million tokens.$0.60$3.60256KTogether AI pricing ↗2026-08-05
CerebrasGPT OSS 120BCerebras lists about 3,000 output tokens per second on its public endpoint.$0.35$0.75131KCerebras hosted pricing ↗2026-08-05
CerebrasGemma 4 31BCerebras lists about 1,850 output tokens per second on its public endpoint.$0.99$1.49131KCerebras hosted pricing ↗2026-08-05
CerebrasGLM 4.7Preview endpoint at about 1,000 output tokens per second. Deprecation is scheduled for August 17, 2026.$2.25$2.75131KCerebras hosted pricing ↗2026-08-05
GroqLlama 3.1 8B InstantGroqCloud production rate at about 560 output tokens per second.$0.05$0.08131KGroq supported models ↗2026-08-05
GroqLlama 3.3 70B VersatileGroqCloud production rate at about 280 output tokens per second.$0.59$0.79131KGroq supported models ↗2026-08-05
GroqGPT OSS 120BGroqCloud production rate at about 500 output tokens per second.$0.15$0.60131KGroq supported models ↗2026-08-05
GroqGPT OSS 20BGroqCloud production rate at about 1,000 output tokens per second.$0.08$0.30131KGroq supported models ↗2026-08-05
GroqQwen 3.6 27BGroqCloud preview rate at about 500 output tokens per second.$0.60$3.00131KGroq supported models ↗2026-08-05
FireworksDeepSeek V4 ProFireworks standard serverless list rate. Cached input costs $0.145 per million tokens.$1.74$3.481MFireworks serverless pricing ↗2026-08-05
FireworksKimi K2.6Fireworks standard serverless list rate. Cached input costs $0.16 per million tokens.$0.95$4.00256KFireworks serverless pricing ↗2026-08-05
FireworksMiniMax M2.7Fireworks standard serverless list rate. Cached input costs $0.06 per million tokens.$0.30$1.20196KFireworks serverless pricing ↗2026-08-05
FireworksGLM 5.1Fireworks standard serverless list rate. Cached input costs $0.26 per million tokens.$1.40$4.40202KFireworks serverless pricing ↗2026-08-05
FireworksGPT OSS 120BFireworks standard serverless list rate. Cached input costs $0.015 per million tokens.$0.15$0.60128KFireworks serverless pricing ↗2026-08-05
MetaLlama 4 MaverickGroqCloud launch rate. The current Groq production catalog omits this endpoint.$0.50$0.771MGroq Llama 4 pricing ↗2026-08-05
MetaLlama 4 ScoutGroqCloud launch rate. The current Groq production catalog omits this endpoint.$0.11$0.3410MGroq Llama 4 pricing ↗2026-08-05
MiniMaxMiniMax M2.5Bedrock standard rates in primary US regions.$0.30$1.20196KAmazon Bedrock pricing ↗2026-08-05
BaiduERNIE 5.0International Qianfan pay-as-you-go rates.$1.40$5.60128KBaidu AI Cloud international pricing ↗2026-08-05
PerplexitySonarSonar API token rates exclude the separate search-context request fee.$1.00$1.00128KPerplexity Sonar API pricing ↗2026-08-05
AI21 LabsJamba 1.5 LargeBedrock standard on-demand rates in US East.$2.00$8.00256KAmazon Bedrock pricing ↗2026-08-05
AI21 LabsJamba 1.5 MiniBedrock standard on-demand rates in US East.$0.20$0.40256KAmazon Bedrock pricing ↗2026-08-05
AnthropicClaude Opus 4.8Standard global Claude API rates.$5.00$25.001MAnthropic API pricing ↗2026-08-05
Sakana AIFugu UltraStandard-context rate; above 272K tokens the rate rises to $10 input and $45 output. Cached input costs $0.50 per million. Base Fugu bills at the rate of whichever model it routes to.$5.00$30.001MSakana pricing ↗2026-08-05
Sakana AIFugu CyberStandard-context rate; above 272K tokens the rate rises to $12 input and $54 output.$6.00$36.001MSakana pricing ↗2026-08-05
Sakana AINamazuJapanese-specialized model built on Kimi K2.6. Cached input costs $0.15 per million tokens.$0.95$4.00256KSakana pricing ↗2026-08-05

17 components

Hardware

A leading ~ marks street estimates and planning allowances. Use those for budget envelopes, and check a live listing when you reach checkout.

Hardware component prices with source and check date
CategoryItemPriceKey specSourceChecked
System24 GB value workstation (Used RTX 3090 class)Whole-system planning allowance based on a dated component budget.~$1,50024 GB VRAMPlanning assumption ↗2026-08-05
GPUGeForce RTX 4090 Founders EditionNVIDIA list starting price for the card alone.$1,59924 GB GDDR6XNVIDIA RTX 4090 product page ↗2026-08-05
GPUGeForce RTX 5090 Founders EditionNVIDIA list starting price for the card alone.$1,99932 GB GDDR7 · 1,792 GB/sNVIDIA RTX 5090 product page ↗2026-08-05
System32 GB performance workstation (RTX 5090 class build)Whole-system planning allowance around a 5090-class build.~$3,20032 GB VRAM · 1,792 GB/sNVIDIA RTX 5090 specifications ↗2026-08-05
SystemDGX Spark Founders EditionNVIDIA marketplace list price for the Founders Edition bundle.$4,699128 GB unified · 273 GB/sNVIDIA DGX Spark marketplace ↗2026-08-05
SystemCapacity + speed local lab (DGX Spark + RTX 5090)Combined planning allowance for separately priced components.~$7,900128 GB unified + 32 GB VRAMNVIDIA DGX Spark hardware ↗2026-08-05
SystemMac mini (M4 Pro, 24 GB)Apple U.S. starting price for the M4 Pro configuration.$1,39924 GB unified · 273 GB/sApple Mac mini newsroom ↗2026-08-05
SystemMac Studio (M4 Max, 36 GB)Apple U.S. starting price for the base M4 Max Studio.$1,99936 GB unified · 410 GB/sApple Mac Studio newsroom ↗2026-08-05
SystemMac Studio (M3 Ultra, 96 GB)Apple U.S. starting price for the M3 Ultra configuration.$3,99996 GB unified · 819 GB/sApple Mac Studio newsroom ↗2026-08-05
GPUUsed GeForce RTX 4090 Founders EditionRounded snapshot of recent used listings with seller and condition variance.~$2,20024 GB GDDR6XeBay used RTX 4090 listings ↗2026-08-05
GPUGeForce RTX 5080 Founders EditionNVIDIA launch list price for the card alone.$99916 GB GDDR7 · 960 GB/sNVIDIA RTX 50 Series announcement ↗2026-08-05
GPUGeForce RTX 5070 TiNVIDIA launch list price for the card alone.$74916 GB GDDR7 · 896 GB/sNVIDIA RTX 50 Series announcement ↗2026-08-05
GPUNVIDIA RTX PRO 6000 Blackwell Workstation EditionNVIDIA Marketplace list price for the workstation card.$13,25096 GB GDDR7 ECC · 1,792 GB/sNVIDIA RTX PRO 6000 marketplace ↗2026-08-05
SystemFramework Desktop Ryzen AI Max+ 395 128 GBFramework launch list price for the 128 GB DIY Edition.$1,999128 GB LPDDR5X · 256 GB/sFramework Desktop announcement ↗2026-08-05
SystemMinisforum MS-S1 MAX 128 GBMinisforum regular list price for the 128 GB and 2 TB configuration.$4,549128 GB LPDDR5X · 256 GB/sMinisforum MS-S1 MAX product page ↗2026-08-05
System16-inch MacBook Pro M4 Max 48 GBApple U.S. launch list price for the 16-core CPU and 40-core GPU configuration.$3,99948 GB unified memory · 546 GB/sApple 16-inch MacBook Pro technical specifications ↗2026-08-05
SystemNVIDIA Jetson AGX Thor Developer KitCurrent NVIDIA developer kit MSRP.$5,499128 GB LPDDR5X · 273 GB/sNVIDIA Jetson product pricing FAQ ↗2026-08-05

14 models and tiers

Video and image generation

These models are priced per generated output, such as a video second or image, instead of per token.

Video and image generation prices with source and check date
KindProviderModelPriceUnitSourceChecked
VideoByteDanceSeedance 2.0BytePlus Dramagic all-purpose rate for generation without video input. The 720p rate is $0.23 per second.$0.56per second of 1080p videoBytePlus Dramagic pricing2026-08-05
VideoByteDanceSeedance 2.5Dreamina offers daily trial credits. A cash API rate has not been published.Unpublishedpublic API rateDreamina Seedance 2.52026-08-05
VideoMiniMaxMiniMax H3MiniMax pay-as-you-go rate. The 2K rate is $0.13 per second.$0.08per second of 768p videoMiniMax API pricing2026-08-05
VideoGoogleVeo 3.1Vertex AI rate for 720p or 1080p video with synchronized audio.$0.40per second with audioGoogle Vertex AI pricing2026-08-05
VideoOpenAISora 2API rate for 720x1280 portrait or 1280x720 landscape video with audio.$0.10per secondOpenAI Sora 2 model pricing2026-08-05
VideoOpenAISora 2 ProAPI rate for 720p output. Higher-resolution output carries a higher rate.$0.30per second of 720p videoOpenAI Sora 2 model pricing2026-08-05
VideoKling via falKling Video 3.0 Standardfal hosted rate. Audio raises the rate to $0.126 per second.$0.08per second without audiofal Kling 3.0 pricing2026-08-05
VideoRunwayGen-4.5Runway Dev charges 12 credits per second at $0.01 per credit. Output is 720p.$0.12per secondRunway API pricing2026-08-05
VideoAlibabaWan 2.6 Image to VideoGlobal deployment rate for video with audio.$0.09per second of 720p videoAlibaba Cloud Model Studio pricing2026-08-05
VideoAlibabaWan 2.6 Image to VideoGlobal deployment rate for video with audio.$0.14per second of 1080p videoAlibaba Cloud Model Studio pricing2026-08-05
ImageBlack Forest LabsFLUX 3The BFL API pricing page has no FLUX 3 endpoint or cash rate.Unpublishedpublic API rateBlack Forest Labs API pricing2026-08-05
ImageIdeogramIdeogram 4.0 DefaultFlat API rate for Generate, Remix, Edit, Reframe, and Replace Background.$0.06per output imageIdeogram API pricing2026-08-05
ImageRecraftRecraft V4.1Standard raster API rate.$0.04per 1 MP raster imageRecraft API pricing2026-08-05
ImageRecraftRecraft V4.1 ProHigh-resolution raster API rate for print and large-format output.$0.25per 4 MP raster imageRecraft API pricing2026-08-05

Next step

Turn the numbers into a stack.

Use these rates as evidence while you plan. The planner turns your workloads, privacy needs, and budget into an explainable local and hosted mix.

Back to the planner →