Together AI Pricing 2026:
Open-Source LLMs at Cloud Scale
Together AI offers pay-per-token access to open-weight models such as Qwen, DeepSeek, GLM, gpt-oss, Gemma and Llama in 2026. Complete pricing guide, fine-tuning costs, and when Together beats AWS or Hugging Face.
Together AI Model Pricing 2026
| Model | Input (per 1M) | Output (per 1M) | Notes |
|---|---|---|---|
| Qwen3.8 Flash | $0.09 | $0.28 | Cheapest paid general model on the list |
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 | Very low output price |
| Llama 3 8B Instruct Lite | $0.14 | $0.14 | Small Llama, flat in/out price |
| gpt-oss-120B | $0.15 | $0.60 | OpenAI open-weight model |
| GLM-5.3-Flash | $0.15 | $0.50 | Fast GLM tier |
| Qwen3 235B A22B Instruct 2507 FP8 | $0.20 | $0.60 | Large MoE at a low price |
| DeepSeek V4.1 Flash | $0.30 | $1.20 | Newer DeepSeek Flash |
| Gemma 4 31B | $0.39 | $0.97 | Google open-weight model |
| Llama 3.3 70B | $1.04 | $1.04 | Last Llama 70B on the serverless list |
| DeepSeek V4 Pro 0813 | $1.32 | $3.96 | Stronger DeepSeek tier |
| Kimi K3 | $3.00 | $15.00 | Premium open-weight model |
Serverless prices from together.ai/pricing, checked 1 October 2026. Together's model list changes often — the Llama 3.1, Mistral 7B, Mixtral and Qwen2.5 72B endpoints listed here in April 2026 are no longer on its serverless price list.
Together AI vs AWS Bedrock vs Groq
A like-for-like price table no longer works: in October 2026 the three providers host different model generations, so the same model is rarely available on all three. Compare per task instead — pick the cheapest model on each platform that passes your quality test, then compare those prices.
Together vs AWS Bedrock: Bedrock is the choice when you need AWS-native IAM, VPC and billing; Together usually offers a wider catalogue of open-weight models at lower per-token prices.
Together vs Groq: Groq's LPU hardware is built for very high tokens-per-second; Together has more model variety. For latency-sensitive apps, test Groq. For variety and batch workloads, Together is usually the better fit.
Together AI Fine-Tuning Pricing
Together AI supports fine-tuning on open-weight models — one of its key differentiators. As of October 2026 its pricing page lists fine-tuning from $0.34 to $100 per 1M training tokens, depending on model size and method (supervised fine-tuning vs preference optimisation). The per-model figures below are as last verified in April 2026:
| Model | Training (per 1M tokens) | Inference after fine-tune |
|---|---|---|
| Llama 3.1 8B | $0.30 | $0.18/M (3× base) |
| Llama 3.1 70B | $3.00 | $1.62/M (3× base) |
| Mistral 7B | $0.30 | $0.30/M |
Fine-tuning Llama 3.1 8B on 5M tokens (April 2026 price): $1.50 total — far below what closed-model fine-tuning cost (OpenAI charged $3/M training tokens for GPT-4o mini, or $15 for this dataset, when last verified in April 2026).
Real-World Cost Example: Moving from GPT-6.1 Sol to Together
Content Generation App (1M input + 1M output tokens/month)
- Current: GPT-6.1 Sol @ $2/M input + $10/M output = ~$12.00/month
- Together Llama 3.3 70B @ $1.04/M (same in/out) × 2M tokens = $2.08/month
- Savings: 83% reduction — if quality is acceptable. A cheaper model such as gpt-oss-120B ($0.15/$0.60) would cost $0.75/month.
- Quality test: run 100 side-by-side comparisons before switching
Together AI Free Tier
- New-account credits change over time and are not listed on the pricing page — check Together's signup page for the current offer
- API compatible with OpenAI SDK (drop-in replacement)
Compare Together AI vs OpenAI Costs
See exactly how much you save by switching to open-source LLMs.
AI Cost Calculator