Together.ai API Pricing (July 2026) — DeepSeek, Qwen, GPT-OSS, Kimi and more

Last updated:

How much does Together AI cost? Together AI API pricing spans $0.05 to $9.00 per million tokens across its open-model catalog. The current flagship is DeepSeek V3.1 at $0.60 / $1.70 per 1M; reasoning workloads run on DeepSeek V4 Pro at $2.10 / $4.40 or DeepSeek R1 at $3.00 / $7.00. The cheapest hosted model is GPT-OSS 20B at $0.05 / $0.20 per 1M, with GPT-OSS 120B at $0.15 / $0.60. New 2026 additions include Kimi K2.6 ($1.20 / $4.50), GLM 5.2 ($1.40 / $4.40), open-weight Qwen 3 routes, and MiniMax M2.7 ($0.30 / $1.20). Together AI is a neutral open-model host with LoRA fine-tuning, dedicated deployments, and an OpenAI-compatible SDK.

How much does each Together AI model cost per million tokens?

Showing 23 current models. .
Showing 23 grouped models from 23 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

Last synced:

Comparing open-model hosts? Novita is a managed multi-model option to quote alongside Together AI, not a replacement for checking the exact model rate below. For GPU rental math, see the self-hosting break-even guide.

Affiliate disclosure: this sponsored link may earn us a commission. It does not affect Together AI table order or pricing data.

Qwen3.5-9B fine-tuning case study

Together's official serverless catalog lists the Qwen3.5-9B base model at the rate shown in the live table above. A separate Fermisense experiment trained a catalog-review specialist from that base and reported a narrow-task win over five frontier configurations. The published derivative has weights but no hosted API price, so it is not a new Together or canonical pricing row. Read the benchmark and cost analysis →

Key facts about Together AI pricing

  • DeepSeek V3.1 — current flagship at $0.60 / $1.70 per 1M tokens (replaces legacy V3 at $1.25 flat).
  • DeepSeek V4 Pro — 2026 reasoning-grade release at $2.10 / $4.40 per 1M tokens.
  • DeepSeek R1 reasoning model at $3.00 / $7.00 per 1M tokens.
  • GPT-OSS 20B — cheapest hosted model at $0.05 / $0.20 per 1M tokens.
  • GPT-OSS 120B at $0.15 / $0.60 per 1M tokens — premium open-weight at sub-Llama cost.
  • Kimi K2.6 long-context model at $1.20 / $4.50 per 1M tokens.
  • GLM 5.2 at $1.40 / $4.40 per 1M tokens; open-weight Qwen 3 routes remain hosted here, while Qwen Plus/Max closed models are first-party Alibaba rows.
  • MiniMax M2.7 at $0.30 / $1.20 via Together, with MiniMax M3 also tracked as a first-party MiniMax row.
  • Llama 3.3 70B still available at $0.88 / $0.88 per 1M tokens — multiple times cheaper than frontier closed models on output.
  • $5 free signup credit; pay-as-you-go beyond, no monthly minimums.
  • LoRA fine-tuning supported on Llama, Mistral, Qwen and DeepSeek families.

Why choose Together AI over Groq or Fireworks?

Together AI's value is breadth, fine-tuning, and production features. The catalog covers Llama (all sizes), Mistral, Qwen, DeepSeek, plus long-tail specialty models that Groq and Fireworks don't carry. Pricing is competitive: Llama 3.3 70B at $0.88 / $0.88 per 1M matches Fireworks within a cent, and DeepSeek V3 at $1.25 / $1.25 is about 2x more than DeepSeek's own API but gives you a US-based host with consistent latency.

LoRA fine-tuning is the killer feature. Together AI supports fine-tuning on every major Llama, Mistral, and Qwen size including the 405B flagship, at roughly $8-$12 per million training tokens. Inference on fine-tuned adapters runs at standard rates plus a small LoRA overhead, which is dramatically cheaper than dedicated deployments. For teams iterating on custom behavior, this is the cleanest path.

Production features include dedicated deployments (guaranteed throughput + lower latency), BYOC options on AWS, and an OpenAI-compatible SDK so migration from OpenAI is usually a base-URL change.

When Together is not the right choice

Groq is 5-10x faster on the same Llama models — if latency is the critical path, Groq wins. Fireworks is slightly cheaper on Llama 3.1 405B. And for frontier reasoning and code quality, GPT-5.4 and Claude 3.5 still lead by a clear margin.

Price History

Only models with a recorded price change are charted here.

DeepSeek V4 Pro

Llama 3.3 70B

Price history tracking started April 2026. Flat model charts stay hidden until a price change is detected.
View pricing changelog →

Frequently asked questions

How much does Together AI charge per token?

Together AI pricing spans $0.05 to $9.00 per million tokens across its open-model catalog. The cheapest model is GPT-OSS 20B at $0.05 / $0.20 per 1M. The current DeepSeek flagship V3.1 is $0.60 / $1.70 per 1M; DeepSeek V4 Pro is $2.10 / $4.40; DeepSeek R1 is $3.00 / $7.00. Llama 3.3 70B is $0.88 flat per 1M. New 2026 additions like Kimi K2.6 ($1.20 / $4.50), GLM 5.2 ($1.40 / $4.40), open-weight Qwen 3 rows, and MiniMax M2.7 ($0.30 / $1.20) are all available. Alibaba Qwen Plus/Max closed models are tracked separately as first-party Alibaba rows. Together AI uses flat per-token pricing with no cached-input discount.

Does Together AI have a free tier?

Yes. Together AI offers $5 in free credits on signup (typically enough for several million tokens of prototyping on smaller Llama or Qwen models). Beyond that, pricing is pay-as-you-go per token with no monthly minimums. Dedicated deployments have their own pricing tier.

How does Together compare to Fireworks?

Together AI and Fireworks are the two main "neutral" open-model hosts and price within a few cents of each other on most SKUs. Llama 3.3 70B is $0.88 on Together vs $0.90 on Fireworks. Fireworks is slightly cheaper on Llama 3.1 405B ($3.00 vs $3.50). Together generally has a broader catalog of smaller and specialty models; Fireworks emphasizes inference speed and function calling.

What's Together's cheapest Llama option?

Llama 3.1 8B is Together AI's cheapest Llama at $0.18 per million input and output tokens. For the 70B tier, Llama 3.3 70B at $0.88 / $0.88 per 1M is the best cost/quality. Mixtral 8x22B at $1.20 / $1.20 is a common sparse MoE alternative if you want lower per-token cost at higher quality than dense 70B.

Does Together support fine-tuning?

Yes. Together AI supports LoRA fine-tuning on most Llama, Mistral, and Qwen models including Llama 3.1 405B. Typical training cost is around $8-$12 per million training tokens depending on model size, with hosted serving of the fine-tuned model at standard rates plus a small LoRA overhead.

What's DeepSeek R1's context window on Together?

DeepSeek R1 on Together AI supports a 64K token context window at $3.00 / $7.00 per 1M tokens. DeepSeek V3 ships with 128K context at $1.25 flat per 1M. If you need longer context, DeepSeek's own API sometimes ships extended context on preview endpoints cheaper than Together AI.

Methodology

Pricing sourced from https://api.together.ai/models on . All prices expressed in USD per 1 million tokens. We track 33 Together AI models spanning DeepSeek (V3.1, V4 Pro, R1), GPT-OSS (20B, 120B), Llama 3.x, Qwen 3.6, Kimi K2.6, GLM 5.1, MiniMax, Mixtral and Mistral catalogs.