LLM Cost Calc

LLM API Pricing — All Models

Per-token pricing for 17 popular language model APIs, in USD per million tokens. Use the cost calculator to price a specific request or your monthly budget.

OpenAI pricing

ModelInput $/1M tokensOutput $/1M tokensContext windowUpdated
GPT-5.6 Terra$2.00$12.00400,000 tok2026-08-08
Standard rate. Long-context requests bill at 2x input / 1.5x output.
GPT-5.6 Sol$5.00$30.00400,000 tok2026-08-08
Standard rate. Long-context requests bill at 2x input / 1.5x output.
GPT-5.6 Luna$0.20$1.20400,000 tok2026-08-08
Standard rate. Long-context requests bill at 2x input / 1.5x output.
GPT-4o$2.50$10.00128,000 tok2026-08-08
Legacy model, still available alongside the GPT-5.6 family.
GPT-4o mini$0.15$0.60128,000 tok2026-08-08
o3$2.00$8.00200,000 tok2026-08-08

Anthropic pricing

ModelInput $/1M tokensOutput $/1M tokensContext windowUpdated
Claude Fable 5$10.00$50.001,000,000 tok2026-08-08
Claude Opus 5$5.00$25.001,000,000 tok2026-08-08
Claude Sonnet 5$2.00$10.001,000,000 tok2026-08-08
Introductory price through 2026-08-31. Rises to $3 / $15 per 1M tokens on 2026-09-01.
Claude Haiku 4.5$1.00$5.00200,000 tok2026-08-08

Google pricing

ModelInput $/1M tokensOutput $/1M tokensContext windowUpdated
Gemini 3.1 Pro$2.00$12.001,048,576 tok2026-08-08
Rate for prompts ≤200K tokens. Above 200K, price doubles to $4 / $18 per 1M tokens.
Gemini 3.6 Flash$1.50$7.501,048,576 tok2026-08-08
Gemini 3.5 Flash-Lite$0.30$2.501,048,576 tok2026-08-08
Gemini 2.5 Flash-Lite$0.10$0.401,048,576 tok2026-08-08
Google is retiring this model on 2026-10-16.

DeepSeek pricing

ModelInput $/1M tokensOutput $/1M tokensContext windowUpdated
DeepSeek V4 Flash$0.14$0.28128,000 tok2026-08-08
Cache-hit input is $0.0028/1M. DeepSeek announced a price increase on 2026-08-06; new rates not yet published — verify before relying on this.
DeepSeek V4 Pro$0.43$0.87128,000 tok2026-08-08
Cache-hit input is $0.003625/1M. DeepSeek announced a price increase on 2026-08-06; new rates not yet published — verify before relying on this.

Groq pricing

ModelInput $/1M tokensOutput $/1M tokensContext windowUpdated
Llama 4 Scout$0.11$0.34131,072 tok2026-08-08
Meta doesn't sell API access directly — this is Groq's hosted rate. Together AI, Fireworks, and Bedrock price Llama differently.

How LLM API pricing works

Most LLM providers charge separately for input tokens (the text you send) and output tokens (the text the model generates), usually priced per million tokens. Output tokens are typically 3-5x more expensive than input tokens since generation is more compute-intensive than reading context. Some providers also offer cached or batch pricing at a discount — this page lists standard, non-cached rates.

Prices are entered manually and may lag behind provider changes. Always confirm current rates on the provider's official pricing page before making budget decisions.