25 popular paid models · updated 2026-08-23 · cache-heavy agentic profile
Data last refreshed 2026-08-28 · refreshes automatically every week
The 25 most-used paid cloud models, repriced on a cache-heavy agentic workload (the pattern typical of coding agents and long-running tools, where ~96% of tokens are cached re-reads). Two columns carry the decision: Cached-in — the rate for repeated tokens — and Eff $/M, the blended cost per million on this profile. Models with a recent price change show the old rate struck out, the new rate in green if it dropped or red if it rose.
| Model | Provider | In $/M | Out $/M | Cached-in $/M | Eff $/M | Score |
|---|---|---|---|---|---|---|
| Opus 4.8 | Anthropic | 15.0 | 75.0 | 1.5 | 2.35 | |
| GPT-5.6 Sol↓ lower | OpenAI | 5.0 4.0 | 30.0 20.0 | 0.5 0.4 | 0.59 | |
| Grok 4.6 | xAI | 2.0 | 6.0 | 0.5 | 0.57 | |
| Kimi K3 | Moonshot (Kimi) | 2.9 | 14.0 | 0.29* | 0.45 | |
| Command R+ | Cohere | 3.0 | 15.0 | 0.3* | 0.44 | |
| Grok 4 | xAI | 3.0 | 15.0 | 0.3* | 0.44 | |
| GLM-5.2 | Zhipu (Z.ai / GLM) | 1.4 | 4.4 | 0.26 | 0.32 | |
| Sonnet 5↑ +50% →2026-09-01 | Anthropic | 2.0 | 10.0 | 0.2 | 0.31 | |
| Gemini 3.1 Pro | 2.0 | 12.0 | 0.2* | 0.30 | ||
| GPT-5.6 Terra | OpenAI | 2.0 | 12.0 | 0.2 | 0.30 | |
| Qwen3.8-Max | Alibaba (Qwen) | 2.0 | 6.0 | 0.17 | 0.27 | |
| Mistral Medium 3.5 | Mistral | 1.5 | 7.5 | 0.15* | 0.22 | |
| Grok 4.3 | xAI | 1.25 | 2.5 | 0.125* | 0.17 | |
| Haiku | Anthropic | 0.8 | 4.0 | 0.08 | 0.13 | |
| Nova Pro | Amazon (Nova) | 0.8 | 3.2 | 0.08* | 0.12 | |
| Llama 3.3 70B | Meta (Llama) | 0.72 | 0.72 | 0.072* | 0.098 | |
| Mistral Large 3 | Mistral | 0.5 | 1.5 | 0.05* | 0.071 | |
| Gemini 3.7 Flash | 0.38 | 1.88 | 0.038* | 0.056 | ||
| GPT-5.6 Luna↓ −80% | OpenAI | 1.0 0.2 | 6.0 1.2 | 0.1 0.02 | 0.030 | |
| Grok 4.1 Fast | xAI | 0.2 | 0.5 | 0.02* | 0.028 | |
| Gemini 3.5 Flash-Lite | 0.15 | 1.25 | 0.015* | 0.024 | ||
| DeepSeek V4-Pro↓ −75% | DeepSeek | 1.74 0.435 | 3.48 0.87 | 0.0145 0.003625 | 0.022 | |
| Mistral Small 4 | Mistral | 0.15 | 0.6 | 0.015* | 0.022 | |
| DeepSeek V4-Flash | DeepSeek | 0.14 | 0.28 | 0.0028 | 0.009 | |
| Nova Lite | Amazon (Nova) | 0.06 | 0.24 | 0.006* | 0.009 |
* cached-in rate estimated at ~10% of input (provider doesn't publish one); the rest are official. The cached-in rate dominates cost on a cache-heavy workload — DeepSeek's is the lowest.Each use case is ranked by the benchmark that governs it — score on the left, effective cost on the right. The row in teal is the best value: the cheapest model that still clears the quality bar for that job. Example weights reflect a dev-heavy workload.
| Use case | Governing score | Example weight |
|---|---|---|
| Production coding | SWE-bench Verified | 75% |
| Content generation | LMArena (preference) | 13% |
| Research & analysis | Reasoning index | 7% |
| Bulk / unattended | Cost-first | fleet |
| Sensitive / private | Local-only | on-device |
| GPT-5.6 Sol | 96.2% | $0.59 |
| Kimi K3 | 93.4%† | $0.43 |
| GPT-5.6 Luna | 93.0%‡ | $0.030 |
| Opus 4.8 | 88.6% | $2.23 |
| Grok 4.6 | 86.6% | $0.57 |
| Gemini 3.1 Pro | 80.6% | $0.30 |
| DeepSeek V4-ProBEST VALUE | 80.6% | $0.022 |
Pick: DeepSeek V4-Pro — 80.6% at ~$0.022/M is the value leader. Keep Sol / Opus for the hardest cross-cutting work.
| Opus 4.8 | top tier | $2.23 |
| Gemini 3.1 Pro | top tier | $0.30 |
| GPT-5.6 Terra | strong | $0.30 |
| Kimi K3 | strong | $0.43 |
| DeepSeek V4-ProBEST VALUE | good | $0.022 |
Pick: Gemini 3.1 Pro or DeepSeek V4-Pro for everyday drafts; Opus only for marquee pieces. (LMArena #1 overall is Claude Fable 5 — the ceiling if writing is critical.)
| GPT-5.6 Sol | frontier | $0.59 |
| Opus 4.8 | frontier | $2.23 |
| Gemini 3.1 Pro | frontier | $0.30 |
| Grok 4.6 | strong | $0.57 |
| DeepSeek V4-ProBEST VALUE | strong | $0.022 |
Pick: Gemini 3.1 Pro (frontier reasoning, cheap) or DeepSeek V4-Pro. Reserve Opus / Sol for high-stakes diligence.
| Nova Lite | light only | $0.009 |
| DeepSeek V4-Flash | good cheap | $0.009 |
| DeepSeek V4-ProBEST VALUE | frontier | $0.022 |
| Gemini 3.5 Flash-Lite | fast light | $0.024 |
| GPT-5.6 Luna | capable | $0.030 |
Pick: DeepSeek V4-Pro — for a few cents more than the cheapest you get frontier quality. Use Luna-batch or Flash-Lite for truly trivial jobs.
| Qwen 4 Coder 32BBEST LOCAL CODER | 32GB Mac | $0 |
| Qwen 3.6-27B | 32GB Mac | $0 |
| qwen2.5:7b | 16–18GB Mac | $0 |
Pick: keeps data on-device. An 8–9B model (qwen2.5:7b) runs on a 16–18GB machine; the strong local coder (Qwen 4 Coder, ~82% SWE) needs 32GB.
As of August 2026 · "Popular" = major paid models across leading providers · figures are API list-price value per 1M tokens. Sources: OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba, Moonshot, Z.ai, Mistral, Amazon, Cohere, Meta published pricing; SWE-bench & LMArena leaderboards. Prices change often — verify before relying on them.