Skip to content
Cost Optimization

OpenAI Batch API:
50% Cost Savings — How It Works in 2026

OpenAI's Batch API cuts costs by 50% across the current GPT-6 models (Astra, Sol and Luna). Here's exactly how it works, when to use it, and how much you can save.

7 min read·Updated October 2026
October 2026 update: OpenAI's current models are GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, and Batch pricing is still 50% below standard — for example GPT-6.1 Sol batch costs $1/$5 and GPT-6 Luna $0.05/$0.25 per 1M tokens. Model prices and examples below use these current models (checked 1 October 2026); earlier versions of this guide used GPT-4o and o-series models, which are now previous-generation models. See our current AI API price comparison.
Batch API vs Standard API — Pricing
50%
discount vs standard
24h
max turnaround time
$1.00
GPT-6.1 Sol input (vs $2.00)
$5.00
GPT-6.1 Sol output (vs $10.00)

Batch API Pricing vs Standard API

ModelStandard InputBatch InputStandard OutputBatch Output
GPT-6.1 Sol$2.00/M$1.00/M$10.00/M$5.00/M
GPT-6 Luna$0.10/M$0.05/M$0.50/M$0.25/M
GPT-6 Astra$10.00/M$5.00/M$50.00/M$25.00/M
text-embedding-3-large (April 2026 price)$0.13/M$0.065/M——

How the Batch API Works

The Batch API processes requests asynchronously — you submit a batch of requests, OpenAI processes them within 24 hours, and you retrieve results when ready:

  1. Create a batch file: JSONL file with up to 50,000 requests
  2. Upload the file: via Files API
  3. Submit the batch: POST to /v1/batches
  4. Poll for completion: GET /v1/batches/{batch_id}
  5. Retrieve results: Download output JSONL file

Key constraints:

  • Maximum 24-hour turnaround (not guaranteed latency)
  • Up to 50,000 requests per batch
  • Enqueued token limits: 90,000 tokens/minute per model by default
  • No streaming — results available only after batch completes

Real-World Batch API Savings

Example: Content Moderation at Scale

Processing 100,000 user reviews per day for sentiment and safety:

  • Average: 200 tokens input + 50 tokens output per review
  • Daily tokens: 20M input + 5M output
  • Standard GPT-6 Luna: 20M × $0.10 + 5M × $0.50 = $2.00 + $2.50 = $4.50/day
  • Batch GPT-6 Luna: 20M × $0.05 + 5M × $0.25 = $1.00 + $1.25 = $2.25/day
  • Annual savings: $821

Example: Research Paper Analysis Pipeline

Processing 10,000 academic papers monthly with GPT-6.1 Sol:

  • Average: 3,000 tokens input + 800 tokens output
  • Monthly tokens: 30M input + 8M output
  • Standard: 30M × $2.00 + 8M × $10.00 = $60 + $80 = $140/month
  • Batch: 30M × $1.00 + 8M × $5.00 = $30 + $40 = $70/month
  • Annual savings: $840

When to Use Batch vs Real-Time API

Use Batch APIUse Standard API
✅ Nightly data processing jobs⚡ User-facing chatbots (response in <5s)
✅ Document classification pipelines⚡ Real-time content generation
✅ Embedding generation for vector DBs⚡ Interactive code completion
✅ Research analysis (offline)⚡ Real-time translation
✅ Content moderation (async)⚡ Customer service with SLA
✅ SEO metadata generation⚡ Safety-critical real-time decisions

Combining Batch API with Prompt Caching

For maximum savings, combine Batch API (50% off) with prompt caching (up to 90% off on cached tokens):

  • Long system prompt (10K tokens) shared across all batch requests
  • After the first request, the shared prompt is billed at the cached-input rate
  • Cached input on GPT-6.1 Sol costs $0.10/M and on GPT-6 Luna $0.01/M — 95% and 90% below their standard input price
  • Check OpenAI's pricing page for how cached input is billed inside batch jobs before budgeting on a stacked discount

Calculate Your Batch API Savings

Enter your monthly token volume and see how much you'd save with Batch API.

AI Cost Calculator