Skip to content

AI API Cost Calculator

Calculate exact monthly costs for GPT-6, Claude Opus 5.5, Gemini 3.x and other AI APIs. Enter your usage patterns and see real-time cost projections.

Select AI Model

Usage Parameters

101K10K100K
505002K8K
502001K4K
Quick Presets

Cost Projection

Model: GPT-6.1 Sol
$90.00
per month
$3.00
Daily
$90.00
Monthly
$1.1K
Annual
Per Request Breakdown
Input cost$0.001000
Output cost$0.002000
Total per request$0.003000
💡 Cost-saving tip
Switch to Gemini 2.5 Flash-Lite and save $86.10/month

AI API Pricing Comparison 2026

Prices are standard (non-batch) rates per 1M tokens. GPT-6 rates apply up to 272K input tokens; longer prompts cost 2× input and 1.5× output. Gemini 3.1 Pro Preview reflects the ≤200k prompt tier. Gemini 3.8 Flash is shown at its promotional rate, which rises to $1.50/$7.50 on January 1, 2027. Last verified: 2026-10-01.

ModelProviderInput /1M tokensOutput /1M tokensContextBest For
GPT-6 AstraOpenAI$10$501.05MOpenAI flagship: hardest reasoning and agentic work
GPT-6.1 SolOpenAI$2$101.05MEveryday coding, analysis and agents
GPT-6 LunaOpenAI$0.1$0.51.05MUltra-high volume, simple tasks
Claude Fable 5.1Anthropic$10$501MAnthropic’s most capable model for the hardest tasks
Claude Opus 5.5Anthropic$4$201MComplex analysis, long agentic tasks
Claude Sonnet 5.5Anthropic$2$101MCoding, long docs, instruction-following
Claude Haiku 4.5Anthropic$1$5200KFast, cost-effective tasks
Gemini 3.1 Pro PreviewGoogle$2$121MComplex multimodal tasks (≤200k prompt tier)
Gemini 3.8 FlashGoogle$0.75$3.751MStrongest Flash model (promo price to Dec 31, 2026)
Gemini 3.5 Flash-LiteGoogle$0.3$2.51MHigh-volume agentic tasks, translation
Gemini 2.5 Flash-LiteGoogle$0.1$0.41MCheapest capable model, max volume
Mistral Medium 3.5Mistral AI$1.5$7.5256KMistral frontier model for agents and coding
Mistral Large 3Mistral AI$0.5$1.5256KEU compliance, multilingual, open-weight
Mistral Small 4Mistral AI$0.15$0.6256KOpen weights, efficient deployment

Frequently Asked Questions

Which AI API is cheapest?

Gemini 2.5 Flash-Lite ($0.10/1M input, $0.40/1M output) and GPT-6 Luna ($0.10/1M input, $0.50/1M output) are currently the cheapest capable production models. Mistral Small 4 at $0.15/$0.60 is the cheapest open-weight option.

How do I reduce AI API costs?

Key strategies: (1) Cache repeated requests, (2) Compress prompts to reduce input tokens, (3) Use faster/cheaper models for simple tasks and only route complex requests to flagship models, (4) Implement request batching, (5) Use prompt caching features offered by Anthropic and OpenAI.

What is a token in AI APIs?

A token is roughly 4 characters or 0.75 words in English. 1,000 tokens ≈ 750 words. A typical paragraph is 100-200 tokens. The word "calculator" is about 2-3 tokens.