LLM API Pricing — All Models
Per-token pricing for 17 popular language model APIs, in USD per million tokens. Use the cost calculator to price a specific request or your monthly budget.
OpenAI pricing
| Model | Input $/1M tokens | Output $/1M tokens | Context window | Updated |
|---|---|---|---|---|
| GPT-5.6 Terra | $2.00 | $12.00 | 400,000 tok | 2026-08-08 |
| Standard rate. Long-context requests bill at 2x input / 1.5x output. | ||||
| GPT-5.6 Sol | $5.00 | $30.00 | 400,000 tok | 2026-08-08 |
| Standard rate. Long-context requests bill at 2x input / 1.5x output. | ||||
| GPT-5.6 Luna | $0.20 | $1.20 | 400,000 tok | 2026-08-08 |
| Standard rate. Long-context requests bill at 2x input / 1.5x output. | ||||
| GPT-4o | $2.50 | $10.00 | 128,000 tok | 2026-08-08 |
| Legacy model, still available alongside the GPT-5.6 family. | ||||
| GPT-4o mini | $0.15 | $0.60 | 128,000 tok | 2026-08-08 |
| o3 | $2.00 | $8.00 | 200,000 tok | 2026-08-08 |
Anthropic pricing
| Model | Input $/1M tokens | Output $/1M tokens | Context window | Updated |
|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | 1,000,000 tok | 2026-08-08 |
| Claude Opus 5 | $5.00 | $25.00 | 1,000,000 tok | 2026-08-08 |
| Claude Sonnet 5 | $2.00 | $10.00 | 1,000,000 tok | 2026-08-08 |
| Introductory price through 2026-08-31. Rises to $3 / $15 per 1M tokens on 2026-09-01. | ||||
| Claude Haiku 4.5 | $1.00 | $5.00 | 200,000 tok | 2026-08-08 |
Google pricing
| Model | Input $/1M tokens | Output $/1M tokens | Context window | Updated |
|---|---|---|---|---|
| Gemini 3.1 Pro | $2.00 | $12.00 | 1,048,576 tok | 2026-08-08 |
| Rate for prompts ≤200K tokens. Above 200K, price doubles to $4 / $18 per 1M tokens. | ||||
| Gemini 3.6 Flash | $1.50 | $7.50 | 1,048,576 tok | 2026-08-08 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1,048,576 tok | 2026-08-08 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1,048,576 tok | 2026-08-08 |
| Google is retiring this model on 2026-10-16. | ||||
DeepSeek pricing
| Model | Input $/1M tokens | Output $/1M tokens | Context window | Updated |
|---|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | 128,000 tok | 2026-08-08 |
| Cache-hit input is $0.0028/1M. DeepSeek announced a price increase on 2026-08-06; new rates not yet published — verify before relying on this. | ||||
| DeepSeek V4 Pro | $0.43 | $0.87 | 128,000 tok | 2026-08-08 |
| Cache-hit input is $0.003625/1M. DeepSeek announced a price increase on 2026-08-06; new rates not yet published — verify before relying on this. | ||||
Groq pricing
| Model | Input $/1M tokens | Output $/1M tokens | Context window | Updated |
|---|---|---|---|---|
| Llama 4 Scout | $0.11 | $0.34 | 131,072 tok | 2026-08-08 |
| Meta doesn't sell API access directly — this is Groq's hosted rate. Together AI, Fireworks, and Bedrock price Llama differently. | ||||
How LLM API pricing works
Most LLM providers charge separately for input tokens (the text you send) and output tokens (the text the model generates), usually priced per million tokens. Output tokens are typically 3-5x more expensive than input tokens since generation is more compute-intensive than reading context. Some providers also offer cached or batch pricing at a discount — this page lists standard, non-cached rates.
Prices are entered manually and may lag behind provider changes. Always confirm current rates on the provider's official pricing page before making budget decisions.