DevTk.AI
AI API PricingGPT-5.6DeepSeek V4Gemini 3.7Grok 4.6Model Pricing

AI API Pricing Comparison (August 2026): Latest Models Side by Side

Updated August 17 API prices for GPT-5.6, Gemini 3.7 Flash, Grok 4.6, DeepSeek V4 peak/off-peak billing, GLM-5.3 status, Kimi K3, Claude, Mistral, and MiniMax.

DevTk.AI 2026-02-19 Updated 2026-08-17 6 min read

The August 17 refresh adds Gemini 3.7 Flash and Grok 4.6, applies Google’s current 2026 introductory Flash price, and replaces DeepSeek’s old flat price with the peak/off-peak schedule that took effect today. GLM-5.3 is also documented, but its pay-as-you-go API and token price are still coming soon, so GLM-5.2 remains the priced Z.AI row.

All USD rows below use standard synchronous rates for cache-miss input and output. Cached-input prices are shown separately. Long-context, Batch, Flex, Fast, regional, and tool charges can change the final bill.

Current Flagship and Balanced API Prices

ModelProviderInputCached inputOutputContextImportant note
GPT-5.6 SolOpenAI$5.00$0.50$30.001.05MLong-context rates above 272K input
GPT-5.6 TerraOpenAI$2.00$0.20$12.001.05MReduced August price
Claude Fable 5Anthropic$10.00$1.00$50.001MCreative model; access restored July 1
Claude Opus 5Anthropic$5.00$0.50$25.001MCurrent Opus flagship
Claude Sonnet 5Anthropic$2.00$0.20$10.001MIntro price through Aug 31
Gemini 3.7 FlashGoogle$0.75$0.075$3.751.05MIntro price through Dec 31; doubles Jan 1
Gemini 3.6 FlashGoogle$0.75$0.075$3.751.05MSame current intro price as 3.7
Gemini 3.1 Pro PreviewGoogle$2.00$0.20$12.001.05M$4/$18 above 200K input
Grok 4.6xAI$2.00$0.50$6.00500K$4/$1/$12 at or above 200K
Grok 4.5xAI$2.00$0.30$6.00500KStill active; cheaper cached input than 4.6
Kimi K3Moonshot AI$3.00$0.30$15.001.05MGlobal API; full weights available
GLM-5.2Z.AI$1.40$0.26$4.401MLatest priced API; GLM-5.3 API price pending
Mistral Medium 3.5Mistral$1.50$0.15$7.50131K90% cached-input discount
Command A+Cohere$2.50-$10.00128KCurrent Cohere flagship

Current Lower-Cost API Prices

ModelProviderInputCached inputOutputContextImportant note
DeepSeek V4 FlashDeepSeek$0.22-$0.44$0.007-$0.014$0.66-$1.321MOff-peak to peak; new schedule effective Aug 17
Xiaomi MiMo-V2.5Xiaomi$0.14$0.0028$0.281MMultimodal; cache writes temporarily free
GPT-5.6 LunaOpenAI$0.20$0.02$1.201.05MReduced 80% from its launch price
Mistral Small 4Mistral$0.15$0.015$0.60131KCurrent rate increased from $0.10/$0.30
MiniMax M3MiniMax$0.30$0.06$1.201M$0.60/$2.40 above 512K input
Gemini 3.5 Flash-LiteGoogle$0.30$0.03$2.501.05MCurrent GA high-volume route
Amazon Nova 2 LiteAmazon$0.30-$2.501MBedrock price may vary by region
Kimi K2.7 CodeMoonshot AI$0.95$0.19$4.00262KGlobal coding model
Kimi K2.6Moonshot AI$0.95$0.16$4.00262KGlobal multimodal model
Claude Haiku 4.5Anthropic$1.00$0.10$5.00200KCurrent low-cost Claude tier
Grok 4.3xAI$1.25$0.20$2.501MLong-context rate above 200K
Jamba Mini 2AI21 Labs$0.20-$0.40256KCurrent jamba-mini alias
Command RCohere$0.50-$1.50128KCorrected from stale $0.15/$0.60 data

China API Prices in CNY

Do not mix these rows into a USD ranking without choosing an exchange rate and region.

ModelInputCached inputOutputContextPrice status
Qwen3.7 PlusÂĨ1.60-ÂĨ6.401MCurrent 20% alias promotion; ÂĨ4.80/ÂĨ19.20 above 256K
Qwen3.7 MaxÂĨ6.00-ÂĨ18.001MCurrent 50% alias promotion
Kimi K2.5ÂĨ4.00ÂĨ0.70ÂĨ21.00262KChina API price

Alibaba’s current pricing page does not list an end date for the Qwen3.7 alias promotions. DevTk.AI records both the current promotional rate and the list price in model notes.

Same Workload Cost Comparison

For 2M cache-miss input tokens plus 500K output tokens, using standard short-context prices:

ModelCost
Xiaomi MiMo-V2.5$0.42
Mistral Small 4$0.60
DeepSeek V4 Flash, off-peak$0.77
GPT-5.6 Luna$1.00
MiniMax M3$1.20
DeepSeek V4 Flash, peak$1.54
Gemini 3.5 Flash-Lite$1.85
Amazon Nova 2 Lite$1.85
Gemini 3.7 Flash, 2026 intro price$3.38
GLM-5.2$5.00
Grok 4.6$7.00
Claude Sonnet 5$9.00
GPT-5.6 Terra$10.00
Kimi K3$13.50
Claude Opus 5$22.50
GPT-5.6 Sol$25.00

This table is a token-cost benchmark, not a quality ranking. Models can use different reasoning-token and output-token amounts for the same task. Tool calls, web search, storage, regional processing, and service tiers are additional.

What Changed This Month

  • OpenAI reduced Terra by 20% and Luna by 80% on both input and output. The former Priority tier is now called Fast mode.
  • DeepSeek’s new peak/off-peak prices are effective. Flash now ranges from $0.22/$0.66 off-peak to $0.44/$1.32 peak for cache-miss input/output.
  • Anthropic now lists Fable 5, Opus 5, and Sonnet 5. Sonnet’s $2/$10 rate ends August 31.
  • Google launched Gemini 3.7 Flash at a $0.75/$3.75 introductory rate and applied the same rate to 3.6 through December 31.
  • xAI launched Grok 4.6 at $2/$6. Its cached input is $0.50, 66.7% above Grok 4.5, and the long-context step begins at 200K.
  • Z.AI announced GLM-5.3, but standard API pricing is not yet published; it is excluded from cost rankings.
  • Mistral Small 4 moved to $0.15/$0.60 and Mistral lists 90% off cached input.

Use the AI Model Pricing Calculator for your own traffic mix. For the market implications, read The August 2026 AI API Price War.

Official Sources

Related Posts