Skip to content
Cost Optimization

OpenAI Batch API Cost 2026:
Save 50% on Every Request

OpenAI's Batch API processes requests asynchronously at exactly 50% off standard pricing. Learn when to use it, how it works, and real-world cost savings for large-scale AI workloads.

10 min read·Updated October 2026
October 2026 update: OpenAI's current models are GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, and Batch pricing is still 50% below standard — for example GPT-6.1 Sol batch costs $1/$5 and GPT-6 Luna $0.05/$0.25 per 1M tokens. Model prices and examples below use these current models (checked 1 October 2026); earlier versions of this guide used GPT-4o and o-series models, which are now previous-generation models. See our current AI API price comparison.
Batch API Savings
50%
discount on all models
24 hrs
maximum turnaround time
50K
max requests per batch
100 MB
max batch file size

OpenAI Batch API Pricing (2026)

ModelStandard InputBatch Input (50% off)Standard OutputBatch Output (50% off)
GPT-6 Astra$10.00/M$5.00/M$50.00/M$25.00/M
GPT-6.1 Sol$2.00/M$1.00/M$10.00/M$5.00/M
GPT-6 Luna$0.10/M$0.05/M$0.50/M$0.25/M
text-embedding-3-large (April 2026 price)$0.13/M$0.065/MN/AN/A

How the Batch API Works

  1. Create a JSONL file with all your requests (one per line)
  2. Upload the file to OpenAI's Files API
  3. Create a batch job referencing the file
  4. Wait for completion (typically 1–6 hours, guaranteed within 24 hours)
  5. Download results from the output file

Real-World Batch API Savings Examples

Content Classification (100,000 items)

  • Each item: 200 input tokens + 20 output tokens = 220 tokens
  • Total: 22M tokens
  • Standard GPT-6 Luna: 20M × $0.10 + 2M × $0.50 = $2 + $1 = $3.00
  • Batch GPT-6 Luna: = $1.50 (saves $1.50)

Document Summarization (10,000 documents)

  • Each document: 2,000 input + 300 output tokens
  • Total: 23M tokens
  • Standard GPT-6.1 Sol: 20M × $2 + 3M × $10 = $40 + $30 = $70
  • Batch GPT-6.1 Sol: = $35 (saves $35)

Product Description Generation (50,000 SKUs)

  • Each product: 150 input + 200 output tokens
  • Standard GPT-6 Luna: 7.5M × $0.10 + 10M × $0.50 = $0.75 + $5 = $5.75
  • Batch GPT-6 Luna: = $2.88 (saves $2.87)

When to Use Batch API vs Real-Time API

Use CaseUse Batch?Reason
Live user chat❌ NoUsers need instant responses
Real-time content moderation❌ NoMust decide before displaying content
Nightly data processing✅ YesResults needed by morning, not instantly
Dataset enrichment✅ YesProcess 1M records overnight
SEO content generation✅ YesGenerate 1,000 articles, no rush
Product catalog analysis✅ YesWeekly processing job
Embedding generation✅ YesOne-time or scheduled vectorization
Sentiment analysis✅ Yes (usually)Dashboard can update daily, not real-time

Anthropic and Google Batch API Equivalents

  • Anthropic Claude: Message Batches API — up to 50% discount, 24-hour processing window
  • Google Gemini: Batch prediction via Vertex AI — pricing varies by model and region
  • AWS Bedrock: Batch inference — 50% discount, similar to OpenAI's offering

Combining Batch API with Other Optimizations

Stack multiple savings techniques for maximum reduction:

  • Batch API (50% off) + Prompt Caching (up to 90% off system prompts) + GPT-6 Luna instead of GPT-6.1 Sol where quality allows (20× cheaper)
  • Combined, these can reduce costs by 95%+ vs naive GPT-6.1 Sol real-time usage

Calculate Your Batch API Savings

See how much you'd save by switching bulk workloads to the Batch API.

AI Cost Calculator