25 popular paid models · updated 2026-08-23 · cache-heavy agentic profile

LLM Cost & Capability

Data last refreshed 2026-08-28 · refreshes automatically every week

The 25 most-used paid cloud models, repriced on a cache-heavy agentic workload (the pattern typical of coding agents and long-running tools, where ~96% of tokens are cached re-reads). Two columns carry the decision: Cached-in — the rate for repeated tokens — and Eff $/M, the blended cost per million on this profile. Models with a recent price change show the old rate struck out, the new rate in green if it dropped or red if it rose.

ModelProviderIn $/MOut $/MCached-in $/MEff $/MScore
Opus 4.8Anthropic15.075.01.52.35
GPT-5.6 Sol↓ lowerOpenAI5.0 4.030.0 20.00.5 0.40.59
Grok 4.6xAI2.06.00.50.57
Kimi K3Moonshot (Kimi)2.914.00.29*0.45
Command R+Cohere3.015.00.3*0.44
Grok 4xAI3.015.00.3*0.44
GLM-5.2Zhipu (Z.ai / GLM)1.44.40.260.32
Sonnet 5↑ +50% →2026-09-01Anthropic2.010.00.20.31
Gemini 3.1 ProGoogle2.012.00.2*0.30
GPT-5.6 TerraOpenAI2.012.00.20.30
Qwen3.8-MaxAlibaba (Qwen)2.06.00.170.27
Mistral Medium 3.5Mistral1.57.50.15*0.22
Grok 4.3xAI1.252.50.125*0.17
HaikuAnthropic0.84.00.080.13
Nova ProAmazon (Nova)0.83.20.08*0.12
Llama 3.3 70BMeta (Llama)0.720.720.072*0.098
Mistral Large 3Mistral0.51.50.05*0.071
Gemini 3.7 FlashGoogle0.381.880.038*0.056
GPT-5.6 Luna↓ −80%OpenAI1.0 0.26.0 1.20.1 0.020.030
Grok 4.1 FastxAI0.20.50.02*0.028
Gemini 3.5 Flash-LiteGoogle0.151.250.015*0.024
DeepSeek V4-Pro↓ −75%DeepSeek1.74 0.4353.48 0.870.0145 0.0036250.022
Mistral Small 4Mistral0.150.60.015*0.022
DeepSeek V4-FlashDeepSeek0.140.280.00280.009
Nova LiteAmazon (Nova)0.060.240.006*0.009

Best model by use case

Each use case is ranked by the benchmark that governs it — score on the left, effective cost on the right. The row in teal is the best value: the cheapest model that still clears the quality bar for that job. Example weights reflect a dev-heavy workload.

Use caseGoverning scoreExample weight
Production codingSWE-bench Verified75%
Content generationLMArena (preference)13%
Research & analysisReasoning index7%
Bulk / unattendedCost-firstfleet
Sensitive / privateLocal-onlyon-device

Production coding

SWE-bench Verified
GPT-5.6 Sol96.2%$0.59
Kimi K393.4%†$0.43
GPT-5.6 Luna93.0%‡$0.030
Opus 4.888.6%$2.23
Grok 4.686.6%$0.57
Gemini 3.1 Pro80.6%$0.30
DeepSeek V4-ProBEST VALUE80.6%$0.022

Pick: DeepSeek V4-Pro — 80.6% at ~$0.022/M is the value leader. Keep Sol / Opus for the hardest cross-cutting work.

Content generation

LMArena preference
Opus 4.8top tier$2.23
Gemini 3.1 Protop tier$0.30
GPT-5.6 Terrastrong$0.30
Kimi K3strong$0.43
DeepSeek V4-ProBEST VALUEgood$0.022

Pick: Gemini 3.1 Pro or DeepSeek V4-Pro for everyday drafts; Opus only for marquee pieces. (LMArena #1 overall is Claude Fable 5 — the ceiling if writing is critical.)

Research & analysis

Reasoning
GPT-5.6 Solfrontier$0.59
Opus 4.8frontier$2.23
Gemini 3.1 Profrontier$0.30
Grok 4.6strong$0.57
DeepSeek V4-ProBEST VALUEstrong$0.022

Pick: Gemini 3.1 Pro (frontier reasoning, cheap) or DeepSeek V4-Pro. Reserve Opus / Sol for high-stakes diligence.

Bulk / unattended fleet

Cost-first
Nova Litelight only$0.009
DeepSeek V4-Flashgood cheap$0.009
DeepSeek V4-ProBEST VALUEfrontier$0.022
Gemini 3.5 Flash-Litefast light$0.024
GPT-5.6 Lunacapable$0.030

Pick: DeepSeek V4-Pro — for a few cents more than the cheapest you get frontier quality. Use Luna-batch or Flash-Lite for truly trivial jobs.

Sensitive / private

Local-only
Qwen 4 Coder 32BBEST LOCAL CODER32GB Mac$0
Qwen 3.6-27B32GB Mac$0
qwen2.5:7b16–18GB Mac$0

Pick: keeps data on-device. An 8–9B model (qwen2.5:7b) runs on a 16–18GB machine; the strong local coder (Qwen 4 Coder, ~82% SWE) needs 32GB.

As of August 2026 · "Popular" = major paid models across leading providers · figures are API list-price value per 1M tokens. Sources: OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba, Moonshot, Z.ai, Mistral, Amazon, Cohere, Meta published pricing; SWE-bench & LMArena leaderboards. Prices change often — verify before relying on them.