Skip to content
API Pricing

Together AI Pricing 2026:
Open-Source LLMs at Cloud Scale

Together AI offers pay-per-token access to open-weight models such as Qwen, DeepSeek, GLM, gpt-oss, Gemma and Llama in 2026. Complete pricing guide, fine-tuning costs, and when Together beats AWS or Hugging Face.

9 min read·Updated October 2026
Together AI Pricing Highlights
$0.09
Qwen3.8 Flash input/1M
$0.14
DeepSeek V4 Flash input/1M
$0.15
gpt-oss-120B input/1M
$1.04
Llama 3.3 70B in/out per 1M

Together AI Model Pricing 2026

ModelInput (per 1M)Output (per 1M)Notes
Qwen3.8 Flash$0.09$0.28Cheapest paid general model on the list
DeepSeek V4 Flash 0731$0.14$0.28Very low output price
Llama 3 8B Instruct Lite$0.14$0.14Small Llama, flat in/out price
gpt-oss-120B$0.15$0.60OpenAI open-weight model
GLM-5.3-Flash$0.15$0.50Fast GLM tier
Qwen3 235B A22B Instruct 2507 FP8$0.20$0.60Large MoE at a low price
DeepSeek V4.1 Flash$0.30$1.20Newer DeepSeek Flash
Gemma 4 31B$0.39$0.97Google open-weight model
Llama 3.3 70B$1.04$1.04Last Llama 70B on the serverless list
DeepSeek V4 Pro 0813$1.32$3.96Stronger DeepSeek tier
Kimi K3$3.00$15.00Premium open-weight model

Serverless prices from together.ai/pricing, checked 1 October 2026. Together's model list changes often — the Llama 3.1, Mistral 7B, Mixtral and Qwen2.5 72B endpoints listed here in April 2026 are no longer on its serverless price list.

Together AI vs AWS Bedrock vs Groq

A like-for-like price table no longer works: in October 2026 the three providers host different model generations, so the same model is rarely available on all three. Compare per task instead — pick the cheapest model on each platform that passes your quality test, then compare those prices.

Together vs AWS Bedrock: Bedrock is the choice when you need AWS-native IAM, VPC and billing; Together usually offers a wider catalogue of open-weight models at lower per-token prices.

Together vs Groq: Groq's LPU hardware is built for very high tokens-per-second; Together has more model variety. For latency-sensitive apps, test Groq. For variety and batch workloads, Together is usually the better fit.

Together AI Fine-Tuning Pricing

Together AI supports fine-tuning on open-weight models — one of its key differentiators. As of October 2026 its pricing page lists fine-tuning from $0.34 to $100 per 1M training tokens, depending on model size and method (supervised fine-tuning vs preference optimisation). The per-model figures below are as last verified in April 2026:

ModelTraining (per 1M tokens)Inference after fine-tune
Llama 3.1 8B$0.30$0.18/M (3× base)
Llama 3.1 70B$3.00$1.62/M (3× base)
Mistral 7B$0.30$0.30/M

Fine-tuning Llama 3.1 8B on 5M tokens (April 2026 price): $1.50 total — far below what closed-model fine-tuning cost (OpenAI charged $3/M training tokens for GPT-4o mini, or $15 for this dataset, when last verified in April 2026).

Real-World Cost Example: Moving from GPT-6.1 Sol to Together

Content Generation App (1M input + 1M output tokens/month)

  • Current: GPT-6.1 Sol @ $2/M input + $10/M output = ~$12.00/month
  • Together Llama 3.3 70B @ $1.04/M (same in/out) × 2M tokens = $2.08/month
  • Savings: 83% reduction — if quality is acceptable. A cheaper model such as gpt-oss-120B ($0.15/$0.60) would cost $0.75/month.
  • Quality test: run 100 side-by-side comparisons before switching

Together AI Free Tier

  • New-account credits change over time and are not listed on the pricing page — check Together's signup page for the current offer
  • API compatible with OpenAI SDK (drop-in replacement)

Compare Together AI vs OpenAI Costs

See exactly how much you save by switching to open-source LLMs.

AI Cost Calculator