AI API Cost Calculator
Calculate exact monthly costs for GPT-6, Claude Opus 5.5, Gemini 3.x and other AI APIs. Enter your usage patterns and see real-time cost projections.
Select AI Model
Usage Parameters
Cost Projection
AI API Pricing Comparison 2026
Prices are standard (non-batch) rates per 1M tokens. GPT-6 rates apply up to 272K input tokens; longer prompts cost 2× input and 1.5× output. Gemini 3.1 Pro Preview reflects the ≤200k prompt tier. Gemini 3.8 Flash is shown at its promotional rate, which rises to $1.50/$7.50 on January 1, 2027. Last verified: 2026-10-01.
| Model | Provider | Input /1M tokens | Output /1M tokens | Context | Best For |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | $10 | $50 | 1.05M | OpenAI flagship: hardest reasoning and agentic work |
| GPT-6.1 Sol | OpenAI | $2 | $10 | 1.05M | Everyday coding, analysis and agents |
| GPT-6 Luna | OpenAI | $0.1 | $0.5 | 1.05M | Ultra-high volume, simple tasks |
| Claude Fable 5.1 | Anthropic | $10 | $50 | 1M | Anthropic’s most capable model for the hardest tasks |
| Claude Opus 5.5 | Anthropic | $4 | $20 | 1M | Complex analysis, long agentic tasks |
| Claude Sonnet 5.5 | Anthropic | $2 | $10 | 1M | Coding, long docs, instruction-following |
| Claude Haiku 4.5 | Anthropic | $1 | $5 | 200K | Fast, cost-effective tasks |
| Gemini 3.1 Pro Preview | $2 | $12 | 1M | Complex multimodal tasks (≤200k prompt tier) | |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1M | Strongest Flash model (promo price to Dec 31, 2026) | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | 1M | High-volume agentic tasks, translation | |
| Gemini 2.5 Flash-Lite | $0.1 | $0.4 | 1M | Cheapest capable model, max volume | |
| Mistral Medium 3.5 | Mistral AI | $1.5 | $7.5 | 256K | Mistral frontier model for agents and coding |
| Mistral Large 3 | Mistral AI | $0.5 | $1.5 | 256K | EU compliance, multilingual, open-weight |
| Mistral Small 4 | Mistral AI | $0.15 | $0.6 | 256K | Open weights, efficient deployment |
Frequently Asked Questions
Which AI API is cheapest?
Gemini 2.5 Flash-Lite ($0.10/1M input, $0.40/1M output) and GPT-6 Luna ($0.10/1M input, $0.50/1M output) are currently the cheapest capable production models. Mistral Small 4 at $0.15/$0.60 is the cheapest open-weight option.
How do I reduce AI API costs?
Key strategies: (1) Cache repeated requests, (2) Compress prompts to reduce input tokens, (3) Use faster/cheaper models for simple tasks and only route complex requests to flagship models, (4) Implement request batching, (5) Use prompt caching features offered by Anthropic and OpenAI.
What is a token in AI APIs?
A token is roughly 4 characters or 0.75 words in English. 1,000 tokens ≈ 750 words. A typical paragraph is 100-200 tokens. The word "calculator" is about 2-3 tokens.