Cost Optimization
OpenAI Batch API Cost 2026:
Save 50% on Every Request
OpenAI's Batch API processes requests asynchronously at exactly 50% off standard pricing. Learn when to use it, how it works, and real-world cost savings for large-scale AI workloads.
10 min read·Updated October 2026
October 2026 update: OpenAI's current models are GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna, and Batch pricing is still 50% below standard — for example GPT-6.1 Sol batch costs $1/$5 and GPT-6 Luna $0.05/$0.25 per 1M tokens. Model prices and examples below use these current models (checked 1 October 2026); earlier versions of this guide used GPT-4o and o-series models, which are now previous-generation models. See our current AI API price comparison.
Batch API Savings
50%
discount on all models
24 hrs
maximum turnaround time
50K
max requests per batch
100 MB
max batch file size
OpenAI Batch API Pricing (2026)
| Model | Standard Input | Batch Input (50% off) | Standard Output | Batch Output (50% off) |
|---|---|---|---|---|
| GPT-6 Astra | $10.00/M | $5.00/M | $50.00/M | $25.00/M |
| GPT-6.1 Sol | $2.00/M | $1.00/M | $10.00/M | $5.00/M |
| GPT-6 Luna | $0.10/M | $0.05/M | $0.50/M | $0.25/M |
| text-embedding-3-large (April 2026 price) | $0.13/M | $0.065/M | N/A | N/A |
How the Batch API Works
- Create a JSONL file with all your requests (one per line)
- Upload the file to OpenAI's Files API
- Create a batch job referencing the file
- Wait for completion (typically 1–6 hours, guaranteed within 24 hours)
- Download results from the output file
Real-World Batch API Savings Examples
Content Classification (100,000 items)
- Each item: 200 input tokens + 20 output tokens = 220 tokens
- Total: 22M tokens
- Standard GPT-6 Luna: 20M × $0.10 + 2M × $0.50 = $2 + $1 = $3.00
- Batch GPT-6 Luna: = $1.50 (saves $1.50)
Document Summarization (10,000 documents)
- Each document: 2,000 input + 300 output tokens
- Total: 23M tokens
- Standard GPT-6.1 Sol: 20M × $2 + 3M × $10 = $40 + $30 = $70
- Batch GPT-6.1 Sol: = $35 (saves $35)
Product Description Generation (50,000 SKUs)
- Each product: 150 input + 200 output tokens
- Standard GPT-6 Luna: 7.5M × $0.10 + 10M × $0.50 = $0.75 + $5 = $5.75
- Batch GPT-6 Luna: = $2.88 (saves $2.87)
When to Use Batch API vs Real-Time API
| Use Case | Use Batch? | Reason |
|---|---|---|
| Live user chat | ❌ No | Users need instant responses |
| Real-time content moderation | ❌ No | Must decide before displaying content |
| Nightly data processing | ✅ Yes | Results needed by morning, not instantly |
| Dataset enrichment | ✅ Yes | Process 1M records overnight |
| SEO content generation | ✅ Yes | Generate 1,000 articles, no rush |
| Product catalog analysis | ✅ Yes | Weekly processing job |
| Embedding generation | ✅ Yes | One-time or scheduled vectorization |
| Sentiment analysis | ✅ Yes (usually) | Dashboard can update daily, not real-time |
Anthropic and Google Batch API Equivalents
- Anthropic Claude: Message Batches API — up to 50% discount, 24-hour processing window
- Google Gemini: Batch prediction via Vertex AI — pricing varies by model and region
- AWS Bedrock: Batch inference — 50% discount, similar to OpenAI's offering
Combining Batch API with Other Optimizations
Stack multiple savings techniques for maximum reduction:
- Batch API (50% off) + Prompt Caching (up to 90% off system prompts) + GPT-6 Luna instead of GPT-6.1 Sol where quality allows (20× cheaper)
- Combined, these can reduce costs by 95%+ vs naive GPT-6.1 Sol real-time usage
Calculate Your Batch API Savings
See how much you'd save by switching bulk workloads to the Batch API.
AI Cost Calculator