Cost Optimization
OpenAI Batch API:
50% Cost Savings — How It Works in 2026
OpenAI's Batch API cuts costs by 50% across the current GPT-6 models (Astra, Sol and Luna). Here's exactly how it works, when to use it, and how much you can save.
7 min read·Updated October 2026
October 2026 update: OpenAI's current models are GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, and Batch pricing is still 50% below standard — for example GPT-6.1 Sol batch costs $1/$5 and GPT-6 Luna $0.05/$0.25 per 1M tokens. Model prices and examples below use these current models (checked 1 October 2026); earlier versions of this guide used GPT-4o and o-series models, which are now previous-generation models. See our current AI API price comparison.
Batch API vs Standard API — Pricing
50%
discount vs standard
24h
max turnaround time
$1.00
GPT-6.1 Sol input (vs $2.00)
$5.00
GPT-6.1 Sol output (vs $10.00)
Batch API Pricing vs Standard API
| Model | Standard Input | Batch Input | Standard Output | Batch Output |
|---|---|---|---|---|
| GPT-6.1 Sol | $2.00/M | $1.00/M | $10.00/M | $5.00/M |
| GPT-6 Luna | $0.10/M | $0.05/M | $0.50/M | $0.25/M |
| GPT-6 Astra | $10.00/M | $5.00/M | $50.00/M | $25.00/M |
| text-embedding-3-large (April 2026 price) | $0.13/M | $0.065/M | — | — |
How the Batch API Works
The Batch API processes requests asynchronously — you submit a batch of requests, OpenAI processes them within 24 hours, and you retrieve results when ready:
- Create a batch file: JSONL file with up to 50,000 requests
- Upload the file: via Files API
- Submit the batch: POST to /v1/batches
- Poll for completion: GET /v1/batches/{batch_id}
- Retrieve results: Download output JSONL file
Key constraints:
- Maximum 24-hour turnaround (not guaranteed latency)
- Up to 50,000 requests per batch
- Enqueued token limits: 90,000 tokens/minute per model by default
- No streaming — results available only after batch completes
Real-World Batch API Savings
Example: Content Moderation at Scale
Processing 100,000 user reviews per day for sentiment and safety:
- Average: 200 tokens input + 50 tokens output per review
- Daily tokens: 20M input + 5M output
- Standard GPT-6 Luna: 20M × $0.10 + 5M × $0.50 = $2.00 + $2.50 = $4.50/day
- Batch GPT-6 Luna: 20M × $0.05 + 5M × $0.25 = $1.00 + $1.25 = $2.25/day
- Annual savings: $821
Example: Research Paper Analysis Pipeline
Processing 10,000 academic papers monthly with GPT-6.1 Sol:
- Average: 3,000 tokens input + 800 tokens output
- Monthly tokens: 30M input + 8M output
- Standard: 30M × $2.00 + 8M × $10.00 = $60 + $80 = $140/month
- Batch: 30M × $1.00 + 8M × $5.00 = $30 + $40 = $70/month
- Annual savings: $840
When to Use Batch vs Real-Time API
| Use Batch API | Use Standard API |
|---|---|
| ✅ Nightly data processing jobs | ⚡ User-facing chatbots (response in <5s) |
| ✅ Document classification pipelines | ⚡ Real-time content generation |
| ✅ Embedding generation for vector DBs | ⚡ Interactive code completion |
| ✅ Research analysis (offline) | ⚡ Real-time translation |
| ✅ Content moderation (async) | ⚡ Customer service with SLA |
| ✅ SEO metadata generation | ⚡ Safety-critical real-time decisions |
Combining Batch API with Prompt Caching
For maximum savings, combine Batch API (50% off) with prompt caching (up to 90% off on cached tokens):
- Long system prompt (10K tokens) shared across all batch requests
- After the first request, the shared prompt is billed at the cached-input rate
- Cached input on GPT-6.1 Sol costs $0.10/M and on GPT-6 Luna $0.01/M — 95% and 90% below their standard input price
- Check OpenAI's pricing page for how cached input is billed inside batch jobs before budgeting on a stacked discount
Calculate Your Batch API Savings
Enter your monthly token volume and see how much you'd save with Batch API.
AI Cost Calculator