Gemini API Pricing 2026:
Gemini 3.8 Flash, Flash-Lite, 3.1 Pro & 2.5
Current Google Gemini API pricing for the Gemini 3.x Flash and Flash-Lite models, Gemini 3.1 Pro Preview and the Gemini 2.5 family, including batch, caching, long-context and free-tier details. Last verified: 2026-10-01.
Google's Gemini lineup moved quickly in 2026. The newest models are the Gemini 3.x Flash family — led by Gemini 3.8 Flash, Google's most capable Flash model — plus the low-cost Flash-Lite models and Gemini 3.1 Pro Preview. The Gemini 2.5 family is still sold and remains the cheapest option at the bottom of the range.
Gemini 3.x Flash and Flash-Lite Models
| Model | Input / 1M tokens | Output / 1M tokens | Batch input / output | Notes |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75* | $3.75* | $0.375 / $1.875* | Most capable Flash model; long-horizon coding and agents |
| Gemini 3.7 Flash | $0.75* | $3.75* | — | Fast, efficient Flash for everyday coding and tool use |
| Gemini 3.6 Flash | $0.75* | $3.75* | — | Previous-generation Flash |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.75 / $4.50 | Earlier Flash model — now pricier than 3.6–3.8 Flash |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.15 / — | High-volume agentic tasks, translation, data processing |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.125 / — | Cheapest Gemini 3.x model (audio input $0.50) |
Gemini 3.1 Pro Preview
| Model | Input / 1M (≤200K) | Input / 1M (>200K) | Output / 1M (≤200K) | Output / 1M (>200K) |
|---|---|---|---|---|
| Gemini 3.1 Pro Preview | $2.00 | $4.00 | $12.00 | $18.00 |
Gemini 2.5 Models (Still Available)
| Model | Input / 1M | Output / 1M | Notes |
|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Cheapest Gemini model; audio input $0.30 |
| Gemini 2.5 Flash | $0.30 | $2.50 | Hybrid reasoning with thinking budgets; audio input $1.00 |
| Gemini 2.5 Pro | $1.25 / $2.50* | $10.00 / $15.00* | *Higher rate for prompts over 200K tokens |
Gemini 2.0 Flash and 2.0 Flash-Lite no longer appear on Google's Gemini API pricing page (checked October 2026). The 2.5 Flash Image model is deprecated and shuts down on October 2, 2026.
Context Caching
Cached input costs about 10% of the standard input price — for example $0.075/M on Gemini 3.8 Flash (promo), $0.03/M on 3.5 Flash-Lite, $0.20/M on 3.1 Pro Preview and $0.01/M on 2.5 Flash-Lite — plus an hourly storage fee ($1.00 per million tokens per hour on most Flash models, $0.50 on 3.8 Flash during the promotion, $4.50 on the Pro models).
Gemini's 1 Million Token Context
Gemini 3.8 Flash, 3.5 Flash-Lite, 3.1 Pro Preview and 2.5 Flash-Lite all accept up to 1,048,576 input tokens. Most frontier models now offer similar windows — Claude Sonnet 5.5 and Opus 5.5 have 1M tokens, the GPT-6 family 1.05M — so the deciding factor is price per token, not context size. Gemini still makes long-context work cheap:
- Full codebase analysis without chunking
- Processing entire legal or financial documents in one request
- Long-form video or audio analysis
- Large dataset inspection without a separate retrieval layer
Gemini Free Tier
The Gemini API free tier covers most current models at no charge, including Gemini 3.8, 3.7, 3.6 and 3.5 Flash, the Flash-Lite models and Gemini 2.5 Pro. Gemini 3.1 Pro Preview is paid-only. Free-tier usage is rate-limited — see Google's rate-limit documentation for current caps — and free-tier prompts may be used to improve Google's products, which the paid tier does not do.
Gemini vs GPT-6 Luna vs Claude Haiku 4.5 (Input Cost)
| Monthly input tokens | Gemini 3.8 Flash* | Gemini 2.5 Flash-Lite | GPT-6 Luna | Claude Haiku 4.5 |
|---|---|---|---|---|
| 10M tokens | $7.50 | $1 | $1 | $10 |
| 100M tokens | $75 | $10 | $10 | $100 |
| 1B tokens | $750 | $100 | $100 | $1,000 |
Input tokens only. *Promotional rate through December 31, 2026. Output prices: 3.8 Flash $3.75, 2.5 Flash-Lite $0.40, GPT-6 Luna $0.50, Haiku 4.5 $5.00 per 1M.
When to Choose Gemini API
- Lowest-cost production option — Gemini 2.5 Flash-Lite ($0.10/$0.40) is the cheapest model from the major providers on output tokens
- Strong mid-tier value — Gemini 3.8 Flash at $0.75/$3.75 (until 2027) costs far less than GPT-6.1 Sol or Claude Sonnet 5.5 ($2/$10)
- Multimodal applications — native image, video and audio input
- Google Cloud ecosystem — Vertex AI, BigQuery and Google Workspace integration
- Free-tier prototyping — the most generous free tier of the major providers
How to Access Gemini API
Get an API key from Google AI Studio (aistudio.google.com) for free-tier usage, or enable billing for the paid tier. For production on Google Cloud, deploy through Vertex AI. No approval process is required for AI Studio API keys.
Frequently Asked Questions
Which Gemini model should new projects start with?
Gemini 3.8 Flash for most workloads, or a Flash-Lite model for high-volume simple tasks. Remember that 3.8 Flash doubles in price on January 1, 2027 ($1.50/$7.50), which still leaves it below GPT-6.1 Sol and Claude Sonnet 5.5.
Which Gemini model is cheapest for production?
Gemini 2.5 Flash-Lite at $0.10/M input and $0.40/M output. Among the Gemini 3.x models, 3.1 Flash-Lite ($0.25/$1.50) is the cheapest.
Should I use Gemini 3.1 Pro Preview in production?
Only with care. Preview models can change, and 3.1 Pro Preview has no free tier. Its $2/$12 price (≤200K tokens) is competitive with GPT-6.1 Sol and Claude Sonnet 5.5 on input but higher on output.
When does Gemini beat OpenAI or Claude on cost?
For high-volume simple tasks, Gemini 2.5 Flash-Lite matches GPT-6 Luna on input ($0.10/M) and is slightly cheaper on output ($0.40 vs $0.50), and both are ten times cheaper than Claude Haiku 4.5. In the middle tier, Gemini 3.8 Flash undercuts the $2/$10 models from OpenAI and Anthropic.
Does Gemini charge more for long prompts?
On the Pro models, yes: prompts over 200K tokens cost more ($4.00/$18.00 on 3.1 Pro Preview, $2.50/$15.00 on 2.5 Pro). The Flash and Flash-Lite models have a single rate.
Calculate Your Gemini API Costs
Compare Gemini vs OpenAI vs Claude for your specific usage volume.
Open API Cost Calculator