Reference prices
Every row cites the provider or product page it came from, with the date we last checked it. Street estimates and planning allowances carry a leading ~ so you can tell them apart from list prices. Confirm against the live page before you spend anything.
Prices last checked August 5, 2026
Open the plannerTokens, translated
Providers price AI in tokens. These translations give the unit some ordinary scale.
The calculator below turns a monthly token guess into an estimated bill.
Estimate a bill
Monthly cost estimator
This is an estimate. The token primer above explains what these volumes mean.
62 models
Input and output rates are USD per million tokens unless a row note says otherwise.
Vercel AI Gateway passes through provider list prices with no markup. Alibaba's current Model Studio pricing and Token Plan allowlist stop at Qwen 3.7, so no Qwen 3.8 price is shown.
| Provider | Model | Input $/M | Output $/M | Context | Source | Checked |
|---|---|---|---|---|---|---|
| DeepSeek | DeepSeek V4 Flash 0731Cache-miss (uncached) rates; cache hits are far cheaper. | $0.14 | $0.28 | 1M | DeepSeek API pricing ↗ | 2026-08-05 |
| DeepSeek | DeepSeek V4 ProCache-miss rates from the DeepSeek pricing table. | $0.44 | $0.87 | 1M | DeepSeek API pricing ↗ | 2026-08-05 |
| OpenAI | GPT-5.6 LunaShort-context standard rates; prompts over 272K are surcharged. | $0.20 | $1.20 | 1M | OpenAI API pricing ↗ | 2026-08-05 |
| OpenAI | GPT-5.6 TerraShort-context standard rates from the OpenAI pricing table. | $2.00 | $12.00 | 1M | OpenAI API pricing ↗ | 2026-08-05 |
| OpenAI | GPT-5.6 SolEscalation tier; short-context standard rates. | $5.00 | $30.00 | 1M | OpenAI API pricing ↗ | 2026-08-05 |
| OpenAI | GPT-5.4 miniStandard rates from the OpenAI pricing table. | $0.75 | $4.50 | 1M | OpenAI API pricing ↗ | 2026-08-05 |
| OpenAI | GPT-5.4 nanoLowest GPT-5.4 text tier on the OpenAI pricing page. | $0.20 | $1.25 | 1M | OpenAI API pricing ↗ | 2026-08-05 |
| Anthropic | Claude Haiku 4.5Standard global Claude API rate. | $1.00 | $5.00 | 200K | Claude API pricing ↗ | 2026-08-05 |
| Anthropic | Claude Sonnet 5Current global Claude API rate. | $2.00 | $10.00 | 1M | Claude API pricing ↗ | 2026-08-05 |
| Anthropic | Claude Opus 5Standard global Claude API rate. | $5.00 | $25.00 | 1M | Claude API pricing ↗ | 2026-08-05 |
| Anthropic | Claude Fable 5Standard rates from the Claude pricing table; batch runs at half price. | $10.00 | $50.00 | 1M | Claude API pricing ↗ | 2026-08-05 |
| Gemini 3.5 Flash-LitePaid-tier standard rates (text/image/video/audio input). | $0.30 | $2.50 | 1M | Gemini Developer API pricing ↗ | 2026-08-05 | |
| Gemini 3.6 FlashPaid-tier standard rates; output includes thinking tokens. | $1.50 | $7.50 | 1M | Gemini Developer API pricing ↗ | 2026-08-05 | |
| Gemini 3.1 ProPaid tier for prompts ≤200K tokens; longer prompts are $4/$18. | $2.00 | $12.00 | 1M | Gemini Developer API pricing ↗ | 2026-08-05 | |
| xAI | Grok 4.3Standard rates for prompts up to 200K tokens. | $1.25 | $2.50 | 1M | xAI API pricing ↗ | 2026-08-05 |
| xAI | Grok 4.5Standard rates for prompts up to 200K tokens. | $2.00 | $6.00 | 500K | xAI API pricing ↗ | 2026-08-05 |
| xAI | Composer 2.5Cursor standard rate. Fast mode costs $3 input and $15 output; xAI also offers the model in Grok Build. | $0.50 | $2.50 | 200K | Cursor Composer 2.5 pricing ↗ | 2026-08-05 |
| Mistral | Mistral Large 3Standard La Plateforme API rates. | $0.50 | $1.50 | 256K | Mistral API pricing ↗ | 2026-08-05 |
| Mistral | Mistral Medium 3.5Standard La Plateforme API rates. | $1.50 | $7.50 | 256K | Mistral API pricing ↗ | 2026-08-05 |
| Mistral | Mistral Small 4Standard La Plateforme API rates. | $0.15 | $0.60 | 256K | Mistral API pricing ↗ | 2026-08-05 |
| Moonshot AI | Kimi K3Cache-miss rates from the official Kimi API table. | $3.00 | $15.00 | 1M | Kimi API pricing ↗ | 2026-08-05 |
| Moonshot AI | Kimi K2.7 CodeCache-miss input rate. Cache hits cost $0.19 per million tokens. | $0.95 | $4.00 | 256K | Kimi K2.7 Code pricing ↗ | 2026-08-05 |
| Moonshot AI | Kimi K2.7 Code HighSpeedCache-miss input rate. Kimi lists about 180 output tokens per second and up to 260 in short contexts. | $1.90 | $8.00 | 256K | Kimi K2.7 Code pricing ↗ | 2026-08-05 |
| Alibaba | Qwen 3.7 MaxSingapore international list rates across the full context window. | $2.50 | $7.50 | 1M | Alibaba Cloud Model Studio pricing ↗ | 2026-08-05 |
| Alibaba | Qwen 3.7 Plus (up to 256K input)Singapore international list rates for prompts up to 256K tokens. | $0.40 | $1.60 | 1M | Alibaba Cloud Model Studio pricing ↗ | 2026-08-05 |
| Alibaba | Qwen 3.7 Plus (256K to 1M input)Singapore international list rates for requests above 256K input tokens. | $1.20 | $4.80 | 1M | Alibaba Cloud Model Studio pricing ↗ | 2026-08-05 |
| Alibaba | Qwen 3.6 FlashSingapore international rates for prompts up to 256K tokens. | $0.25 | $1.50 | 1M | Alibaba Cloud Model Studio pricing ↗ | 2026-08-05 |
| Z.ai | GLM-5Standard cache-miss rates from the Z.ai pricing table. | $1.00 | $3.20 | 200K | Z.ai model pricing ↗ | 2026-08-05 |
| Z.ai | GLM-5.2Standard cache-miss rates from the Z.ai pricing table. | $1.40 | $4.40 | 200K | Z.ai model pricing ↗ | 2026-08-05 |
| Z.ai | GLM-4.5-AirStandard cache-miss rates from the Z.ai pricing table. | $0.20 | $1.10 | 128K | Z.ai model pricing ↗ | 2026-08-05 |
| Amazon | Amazon Nova ProBedrock standard on-demand rates in primary US regions. | $0.80 | $3.20 | 300K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| Amazon | Amazon Nova LiteBedrock standard on-demand rates in primary US regions. | $0.06 | $0.24 | 300K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| Cohere | Command AStandard production API rates. | $2.50 | $10.00 | 256K | Cohere Command A model card ↗ | 2026-08-05 |
| Thinking Machines Lab | Inkling-SmallTinker serverless inference beta rate. Cached input costs $0.06 per million tokens. | $0.30 | $1.20 | 256K | Tinker models and pricing ↗ | 2026-08-05 |
| Thinking Machines Lab | InklingTinker serverless inference beta rate. Cached input costs $0.17 per million tokens. | $1.00 | $4.05 | 256K | Tinker models and pricing ↗ | 2026-08-05 |
| NVIDIA | Nemotron 3 Nano 30B A3BBedrock standard on-demand rate in the primary US regions. | $0.06 | $0.24 | 256K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| NVIDIA | Nemotron 3 Super 120B A12BBedrock standard on-demand rate in the primary US regions. | $0.15 | $0.65 | 256K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| NVIDIA | Nemotron 3 Ultra 550B A55BTogether serverless list rate. Cached input costs $0.20 per million tokens. | $0.60 | $3.60 | 256K | Together AI pricing ↗ | 2026-08-05 |
| Cerebras | GPT OSS 120BCerebras lists about 3,000 output tokens per second on its public endpoint. | $0.35 | $0.75 | 131K | Cerebras hosted pricing ↗ | 2026-08-05 |
| Cerebras | Gemma 4 31BCerebras lists about 1,850 output tokens per second on its public endpoint. | $0.99 | $1.49 | 131K | Cerebras hosted pricing ↗ | 2026-08-05 |
| Cerebras | GLM 4.7Preview endpoint at about 1,000 output tokens per second. Deprecation is scheduled for August 17, 2026. | $2.25 | $2.75 | 131K | Cerebras hosted pricing ↗ | 2026-08-05 |
| Groq | Llama 3.1 8B InstantGroqCloud production rate at about 560 output tokens per second. | $0.05 | $0.08 | 131K | Groq supported models ↗ | 2026-08-05 |
| Groq | Llama 3.3 70B VersatileGroqCloud production rate at about 280 output tokens per second. | $0.59 | $0.79 | 131K | Groq supported models ↗ | 2026-08-05 |
| Groq | GPT OSS 120BGroqCloud production rate at about 500 output tokens per second. | $0.15 | $0.60 | 131K | Groq supported models ↗ | 2026-08-05 |
| Groq | GPT OSS 20BGroqCloud production rate at about 1,000 output tokens per second. | $0.08 | $0.30 | 131K | Groq supported models ↗ | 2026-08-05 |
| Groq | Qwen 3.6 27BGroqCloud preview rate at about 500 output tokens per second. | $0.60 | $3.00 | 131K | Groq supported models ↗ | 2026-08-05 |
| Fireworks | DeepSeek V4 ProFireworks standard serverless list rate. Cached input costs $0.145 per million tokens. | $1.74 | $3.48 | 1M | Fireworks serverless pricing ↗ | 2026-08-05 |
| Fireworks | Kimi K2.6Fireworks standard serverless list rate. Cached input costs $0.16 per million tokens. | $0.95 | $4.00 | 256K | Fireworks serverless pricing ↗ | 2026-08-05 |
| Fireworks | MiniMax M2.7Fireworks standard serverless list rate. Cached input costs $0.06 per million tokens. | $0.30 | $1.20 | 196K | Fireworks serverless pricing ↗ | 2026-08-05 |
| Fireworks | GLM 5.1Fireworks standard serverless list rate. Cached input costs $0.26 per million tokens. | $1.40 | $4.40 | 202K | Fireworks serverless pricing ↗ | 2026-08-05 |
| Fireworks | GPT OSS 120BFireworks standard serverless list rate. Cached input costs $0.015 per million tokens. | $0.15 | $0.60 | 128K | Fireworks serverless pricing ↗ | 2026-08-05 |
| Meta | Llama 4 MaverickGroqCloud launch rate. The current Groq production catalog omits this endpoint. | $0.50 | $0.77 | 1M | Groq Llama 4 pricing ↗ | 2026-08-05 |
| Meta | Llama 4 ScoutGroqCloud launch rate. The current Groq production catalog omits this endpoint. | $0.11 | $0.34 | 10M | Groq Llama 4 pricing ↗ | 2026-08-05 |
| MiniMax | MiniMax M2.5Bedrock standard rates in primary US regions. | $0.30 | $1.20 | 196K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| Baidu | ERNIE 5.0International Qianfan pay-as-you-go rates. | $1.40 | $5.60 | 128K | Baidu AI Cloud international pricing ↗ | 2026-08-05 |
| Perplexity | SonarSonar API token rates exclude the separate search-context request fee. | $1.00 | $1.00 | 128K | Perplexity Sonar API pricing ↗ | 2026-08-05 |
| AI21 Labs | Jamba 1.5 LargeBedrock standard on-demand rates in US East. | $2.00 | $8.00 | 256K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| AI21 Labs | Jamba 1.5 MiniBedrock standard on-demand rates in US East. | $0.20 | $0.40 | 256K | Amazon Bedrock pricing ↗ | 2026-08-05 |
| Anthropic | Claude Opus 4.8Standard global Claude API rates. | $5.00 | $25.00 | 1M | Anthropic API pricing ↗ | 2026-08-05 |
| Sakana AI | Fugu UltraStandard-context rate; above 272K tokens the rate rises to $10 input and $45 output. Cached input costs $0.50 per million. Base Fugu bills at the rate of whichever model it routes to. | $5.00 | $30.00 | 1M | Sakana pricing ↗ | 2026-08-05 |
| Sakana AI | Fugu CyberStandard-context rate; above 272K tokens the rate rises to $12 input and $54 output. | $6.00 | $36.00 | 1M | Sakana pricing ↗ | 2026-08-05 |
| Sakana AI | NamazuJapanese-specialized model built on Kimi K2.6. Cached input costs $0.15 per million tokens. | $0.95 | $4.00 | 256K | Sakana pricing ↗ | 2026-08-05 |
17 components
A leading ~ marks street estimates and planning allowances. Use those for budget envelopes, and check a live listing when you reach checkout.
| Category | Item | Price | Key spec | Source | Checked |
|---|---|---|---|---|---|
| System | 24 GB value workstation (Used RTX 3090 class)Whole-system planning allowance based on a dated component budget. | ~$1,500 | 24 GB VRAM | Planning assumption ↗ | 2026-08-05 |
| GPU | GeForce RTX 4090 Founders EditionNVIDIA list starting price for the card alone. | $1,599 | 24 GB GDDR6X | NVIDIA RTX 4090 product page ↗ | 2026-08-05 |
| GPU | GeForce RTX 5090 Founders EditionNVIDIA list starting price for the card alone. | $1,999 | 32 GB GDDR7 · 1,792 GB/s | NVIDIA RTX 5090 product page ↗ | 2026-08-05 |
| System | 32 GB performance workstation (RTX 5090 class build)Whole-system planning allowance around a 5090-class build. | ~$3,200 | 32 GB VRAM · 1,792 GB/s | NVIDIA RTX 5090 specifications ↗ | 2026-08-05 |
| System | DGX Spark Founders EditionNVIDIA marketplace list price for the Founders Edition bundle. | $4,699 | 128 GB unified · 273 GB/s | NVIDIA DGX Spark marketplace ↗ | 2026-08-05 |
| System | Capacity + speed local lab (DGX Spark + RTX 5090)Combined planning allowance for separately priced components. | ~$7,900 | 128 GB unified + 32 GB VRAM | NVIDIA DGX Spark hardware ↗ | 2026-08-05 |
| System | Mac mini (M4 Pro, 24 GB)Apple U.S. starting price for the M4 Pro configuration. | $1,399 | 24 GB unified · 273 GB/s | Apple Mac mini newsroom ↗ | 2026-08-05 |
| System | Mac Studio (M4 Max, 36 GB)Apple U.S. starting price for the base M4 Max Studio. | $1,999 | 36 GB unified · 410 GB/s | Apple Mac Studio newsroom ↗ | 2026-08-05 |
| System | Mac Studio (M3 Ultra, 96 GB)Apple U.S. starting price for the M3 Ultra configuration. | $3,999 | 96 GB unified · 819 GB/s | Apple Mac Studio newsroom ↗ | 2026-08-05 |
| GPU | Used GeForce RTX 4090 Founders EditionRounded snapshot of recent used listings with seller and condition variance. | ~$2,200 | 24 GB GDDR6X | eBay used RTX 4090 listings ↗ | 2026-08-05 |
| GPU | GeForce RTX 5080 Founders EditionNVIDIA launch list price for the card alone. | $999 | 16 GB GDDR7 · 960 GB/s | NVIDIA RTX 50 Series announcement ↗ | 2026-08-05 |
| GPU | GeForce RTX 5070 TiNVIDIA launch list price for the card alone. | $749 | 16 GB GDDR7 · 896 GB/s | NVIDIA RTX 50 Series announcement ↗ | 2026-08-05 |
| GPU | NVIDIA RTX PRO 6000 Blackwell Workstation EditionNVIDIA Marketplace list price for the workstation card. | $13,250 | 96 GB GDDR7 ECC · 1,792 GB/s | NVIDIA RTX PRO 6000 marketplace ↗ | 2026-08-05 |
| System | Framework Desktop Ryzen AI Max+ 395 128 GBFramework launch list price for the 128 GB DIY Edition. | $1,999 | 128 GB LPDDR5X · 256 GB/s | Framework Desktop announcement ↗ | 2026-08-05 |
| System | Minisforum MS-S1 MAX 128 GBMinisforum regular list price for the 128 GB and 2 TB configuration. | $4,549 | 128 GB LPDDR5X · 256 GB/s | Minisforum MS-S1 MAX product page ↗ | 2026-08-05 |
| System | 16-inch MacBook Pro M4 Max 48 GBApple U.S. launch list price for the 16-core CPU and 40-core GPU configuration. | $3,999 | 48 GB unified memory · 546 GB/s | Apple 16-inch MacBook Pro technical specifications ↗ | 2026-08-05 |
| System | NVIDIA Jetson AGX Thor Developer KitCurrent NVIDIA developer kit MSRP. | $5,499 | 128 GB LPDDR5X · 273 GB/s | NVIDIA Jetson product pricing FAQ ↗ | 2026-08-05 |
14 models and tiers
These models are priced per generated output, such as a video second or image, instead of per token.
| Kind | Provider | Model | Price | Unit | Source | Checked |
|---|---|---|---|---|---|---|
| Video | ByteDance | Seedance 2.0BytePlus Dramagic all-purpose rate for generation without video input. The 720p rate is $0.23 per second. | $0.56 | per second of 1080p video | BytePlus Dramagic pricing | 2026-08-05 |
| Video | ByteDance | Seedance 2.5Dreamina offers daily trial credits. A cash API rate has not been published. | Unpublished | public API rate | Dreamina Seedance 2.5 | 2026-08-05 |
| Video | MiniMax | MiniMax H3MiniMax pay-as-you-go rate. The 2K rate is $0.13 per second. | $0.08 | per second of 768p video | MiniMax API pricing | 2026-08-05 |
| Video | Veo 3.1Vertex AI rate for 720p or 1080p video with synchronized audio. | $0.40 | per second with audio | Google Vertex AI pricing | 2026-08-05 | |
| Video | OpenAI | Sora 2API rate for 720x1280 portrait or 1280x720 landscape video with audio. | $0.10 | per second | OpenAI Sora 2 model pricing | 2026-08-05 |
| Video | OpenAI | Sora 2 ProAPI rate for 720p output. Higher-resolution output carries a higher rate. | $0.30 | per second of 720p video | OpenAI Sora 2 model pricing | 2026-08-05 |
| Video | Kling via fal | Kling Video 3.0 Standardfal hosted rate. Audio raises the rate to $0.126 per second. | $0.08 | per second without audio | fal Kling 3.0 pricing | 2026-08-05 |
| Video | Runway | Gen-4.5Runway Dev charges 12 credits per second at $0.01 per credit. Output is 720p. | $0.12 | per second | Runway API pricing | 2026-08-05 |
| Video | Alibaba | Wan 2.6 Image to VideoGlobal deployment rate for video with audio. | $0.09 | per second of 720p video | Alibaba Cloud Model Studio pricing | 2026-08-05 |
| Video | Alibaba | Wan 2.6 Image to VideoGlobal deployment rate for video with audio. | $0.14 | per second of 1080p video | Alibaba Cloud Model Studio pricing | 2026-08-05 |
| Image | Black Forest Labs | FLUX 3The BFL API pricing page has no FLUX 3 endpoint or cash rate. | Unpublished | public API rate | Black Forest Labs API pricing | 2026-08-05 |
| Image | Ideogram | Ideogram 4.0 DefaultFlat API rate for Generate, Remix, Edit, Reframe, and Replace Background. | $0.06 | per output image | Ideogram API pricing | 2026-08-05 |
| Image | Recraft | Recraft V4.1Standard raster API rate. | $0.04 | per 1 MP raster image | Recraft API pricing | 2026-08-05 |
| Image | Recraft | Recraft V4.1 ProHigh-resolution raster API rate for print and large-format output. | $0.25 | per 4 MP raster image | Recraft API pricing | 2026-08-05 |