Skip to content
Cloud AI Pricing

Google Vertex AI Pricing 2026:
Gemini 2.5 Flash, Pro & Enterprise Costs

Complete Google Vertex AI pricing guide for 2026 — Gemini 2.5 Flash-Lite, Flash, and Pro. How Vertex AI compares to Gemini API direct, compliance features, and when to choose each path. Last verified: 2026-10-01.

10 min read·Updated October 2026
Gemini 2.0 shutdown notice: Gemini 2.0 Flash and Gemini 2.0 Flash-Lite are scheduled for shutdown on 2026-06-01. The current production models on Vertex AI are the Gemini 2.5 family. This page reflects Gemini 2.5 pricing.
October 2026 update: Gemini 2.5 token rates below were re-checked on the Vertex AI pricing page on 1 October 2026 and are unchanged. Vertex AI now also offers the Gemini 3.x models: Gemini 3.8 Flash at an introductory $0.75/$3.75 per 1M tokens through December 31, 2026 (rising to $1.50/$7.50 from January 1, 2027), and Gemini 3.1 Pro Preview at $2/$12 (prompts up to 200K tokens). Fine-tuning and embedding figures are as last verified in April 2026. See our current AI API price comparison.
Vertex AI Gemini 2.5 Pricing at a Glance
$0.10/M
Flash-Lite input (cheapest)
$0.30/M
Gemini 2.5 Flash input
$1.25/M
Gemini 2.5 Pro input
1M tokens
Context window (all tiers)

Gemini 2.5 on Vertex AI — Current Model Pricing

ModelInput / 1M tokensOutput / 1M tokensContext windowBest for
Gemini 2.5 Flash-Lite$0.10$0.401M tokensHigh-volume classification, chatbots, simple tasks
Gemini 2.5 Flash$0.30$2.501M tokensMid-range reasoning, long documents, coding
Gemini 2.5 Pro$1.25$10.001M tokensComplex reasoning, full codebase analysis, research
text-embedding-005$0.025N/A2KSemantic search, RAG ingestion

A key Gemini 2.5 advantage: 1M token context window is available at ALL tiers, including the cheapest Flash-Lite at $0.10/M. OpenAI's current GPT-6 family also offers ~1M context, so context size is no longer the differentiator it was.

Vertex AI vs Gemini API Direct: Key Differences

FeatureVertex AIGemini API (AI Studio)
PricingSame token ratesSame token rates
Free tier$300 Google Cloud creditsGenerous free tier (Flash-Lite)
Enterprise compliance (GDPR, HIPAA, SOC 2)Full supportLimited
Data residencyEU, US, APAC regionsUS primarily
Fine-tuning (supervised)Full supportLimited
Batch predictions (50% off)YesYes
Google Cloud integration (BigQuery, GCS)NativeNot available
Private networking (VPC)VPC Service ControlsNot available

Choose Gemini API direct for development and cost-sensitive production. Choose Vertex AI when you need enterprise compliance, data residency, or deep GCP integration.

Gemini vs OpenAI GPT-6 on Price (October 2026)

TierGoogle modelGoogle input/1MOpenAI modelOpenAI input/1MPrice gap
BudgetGemini 2.5 Flash-Lite$0.10GPT-6 Luna$0.10Same input price
Mid-rangeGemini 3.8 Flash*$0.75GPT-6.1 Sol$2.00Google 2.7× cheaper*
PremiumGemini 3.1 Pro Preview$2.00GPT-6 Astra$10.00Google 5× cheaper

*Gemini 3.8 Flash introductory price through December 31, 2026; from January 1, 2027 it is $1.50/M input, still cheaper than GPT-6.1 Sol. Output prices differ too — see each provider's pricing page.

Real-World Vertex AI Cost Example

Document Processing Pipeline (1M pages/month)

  • Average page: 500 tokens input + 200 tokens output
  • Total: 500M input + 200M output tokens
  • Gemini 2.5 Flash-Lite: $50 + $80 = $130/month
  • Gemini 2.5 Flash: $150 + $500 = $650/month
  • GPT-6 Luna (OpenAI): $50 + $100 = $150/month
  • GPT-6.1 Sol (OpenAI): $1,000 + $2,000 = $3,000/month

Vertex AI Fine-Tuning Costs

  • Gemini 2.5 Flash fine-tuning: $8.00 per 1M training tokens (as last verified April 2026)
  • Fine-tuned model inference: standard Gemini 2.5 Flash pricing applies
  • Minimum training dataset: 100 examples
  • Typical fine-tune: 10,000 examples = ~5M tokens = ~$40 one-time cost

When to Choose Vertex AI

  • You're already on Google Cloud — consolidated billing, existing credits, no new vendor
  • HIPAA/GDPR compliance required — Vertex AI is a Google Cloud HIPAA-eligible service
  • Data needs to stay in EU or APAC — Vertex supports regional data residency
  • You need 1M context at the lowest cost — Flash-Lite at $0.10/M with 1M context; no OpenAI equivalent
  • You want batch processing discounts — 50% off for batch predictions via Vertex

Compare Vertex AI vs Azure vs Direct API

Calculate which cloud AI platform is cheapest for your workload volume.

AI API Cost Calculator