Compare models
Price, context, and benchmarks side by side
Context window, price per million tokens, reasoning levels, and benchmark results with the date they were checked. Unconfirmed fields are marked; provider pricing and capability details can change.
This is the full comparison table for model APIs (GPT, Claude, Gemini, and so on). For launches in the last 28 days, open New models. For Cursor, Claude Code, Muse Code, Lovable and similar products, see Agentic harnesses. Right now that includes 13 recent launches.
Run the model picker4 questions → frontier, balanced, or budget shortlist
Task guides
Choose by the work you need done
Short guides for coding, writing, research, RAG, voice, and more — then compare the models that actually fit.
Task guide · coding
Best AI for coding: choose the right model and coding agent
Alternatives guide · Cursor
Cursor alternatives: compare editor, terminal, and open-source agents
Alternatives guide · open source
Open-source coding agent alternatives: choose the client, model, and privacy boundary
Task guide · writing
Best AI for writing: match quality, voice, and review effort
Task guide · research
Best AI for research: compare context, sources, and verification
Task guide · meetings
Best AI for meeting notes: optimise for capture, privacy, and follow-through
| Model | Tier | Context | Max output | In / 1M | Out / 1M | Weights |
|---|---|---|---|---|---|---|
| GPT-5.6 SolOpenAI | Frontier | 1.05M | 128K | $5 | $30 | Closed |
| GPT-5.6 TerraOpenAI | Balanced | 1.05M | 128K | $2 | $12 | Closed |
| GPT-5.6 LunaOpenAI | Budget | 1.05M | 128K | $0.20 | $1.20 | Closed |
| Claude Opus 5Anthropic | Frontier | 1M | 128K | $5 | $25 | Closed |
| Claude Fable 5Anthropic | Frontier | 1M | 128K | $10 | $50 | Closed |
| Claude Sonnet 5Anthropic | Balanced | 1M | 128K | $2 | $10 | Closed |
| Gemini 3.1 ProGoogle | Frontier | 1.05M | 66K | $2 | $12 | Closed |
| Gemini 3.6 FlashGoogle | Budget | 1.05M | 66K | $0.75 | $3.75 | Closed |
| Gemini 3.7 FlashGoogle | Budget | 1.05M | 66K | $0.75 | $3.75 | Closed |
| Gemini 3.5 Flash-LiteGoogle | Budget | 1.05M | 66K | $0.30 | $2.50 | Closed |
| Kimi K3Moonshot AI | Frontier | 1.05M | 1M | $3 | $15 | Open |
| Kimi K2.6Moonshot AI | Budget | 262K | Not verified | $0.95 | $4 | Open |
| GLM 5.2Z.ai | Frontier | 1M | 128K | $1.40 | $4.40 | Open |
| GLM 5.3Z.ai | Frontier | 1M | 128K | Not verified | Not verified | Closed |
| Grok 4.5SpaceXAI | Balanced | 500K | Not verified | $2 | $6 | Closed |
| Grok 4.6SpaceXAI | Frontier | 500K | Not verified | $2 | $6 | Closed |
| Claude Haiku 4.5Anthropic | Budget | 200K | 64K | $1 | $5 | Closed |
| DeepSeek V4-FlashDeepSeek | Budget | 1M | 384K | $0.14 | $0.28 | Open |
| DeepSeek V4-ProDeepSeek | Balanced | 1M | 384K | $0.435 | $0.87 | Open |
| Ling 3.0 FlashInclusionAI | Budget | 262K | Not verified | $0.075 | $0.22 | Open |
| Qwen3.8-MaxQwen | Frontier | 1M | 131K | $2 | $6 | Closed |
| Qwen3.8-2.4T-A95BQwen | Frontier | 262K | 131K | Not verified | Not verified | Open |
| Qwen3.8-27BQwen | Balanced | 262K | 131K | Not verified | Not verified | Open |
| Muse Spark 1.2Meta | Frontier | 1.05M | Not verified | $1.25 | $4.25 | Closed |
| Muse Glimmer 30BMeta | Balanced | 131K | Not verified | $0 | $0 | Open |
Compare two models
Prices are USD per million tokens at standard rates, excluding batch and caching discounts. Last verified August 14, 2026. Model pricing and capability in this category change frequently — check the provider before committing spend.
Benchmark
DeepSWE 1.1 — coding agent performance and cost
Score alone hides the decision. Cost per completed task sits alongside it, because a model a few points behind at a fraction of the price is usually the better build.
Local leader
Claude Opus 5 [max]
74%
Rows shown
24
Highest published reasoning effort per model (not best Pass@1)
Snapshot date
2026-08-13
Mirrored from deepswe.datacurve.ai
Better is toward the top-right (higher pass rate, lower cost). X-axis is reversed to match DeepSWE’s public chart. v1.1 uses average cost / tokens / steps; v1 uses published medians.
| # | Model | Pass@1 | Cost / task | Tokens / task | Steps / task |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 [max] | 74% | $11.84 | 118k | 99 |
| 2 | GPT-5.6 Sol [max] | 73% | $8.39 | 60k | 61 |
| 3 | Claude Fable 5 [max] | 70% | $21.63 | 119k | 88 |
| 4 | GPT-5.6 Terra [max] | 70% | $4.95 | 72k | 76 |
| 5 | Kimi K3 [max] | 69% | $4.65 | 82k | 98 |
| 6 | GPT-5.6 Luna [max] | 67% | $3.03 | 73k | 102 |
| 7 | GPT-5.5 [xhigh] | 67% | $7.23 | 46k | 82 |
| 8 | Grok 4.6 [xhigh] | 67% | $5.50 | 71k | 87 |
| 9 | Gemini 3.7 Flash [high] | 65% | $2.18 | 107k | 125 |
| 10 | DeepSeek V4-Pro [max] | 63% | $0.24 | 106k | 155 |
| 11 | Claude Opus 4.8 [max] | 59% | $13.22 | 135k | 120 |
| 12 | Qwen3.8-Max [xhigh] | 58% | $3.73 | 95k | 111 |
| 13 | Muse Spark 1.2 [xhigh] | 55% | $3.70 | 99k | 101 |
| 14 | Claude Sonnet 5 [max] | 54% | $26.40 | 214k | 268 |
| 15 | Grok 4.5 [high] | 54% | $2.42 | 36k | 61 |
| 16 | DeepSeek V4-Flash [max] | 53% | $0.10 | 108k | 153 |
| 17 | Muse Spark 1.1 [xhigh] | 53% | $2.36 | 74k | 96 |
| 18 | GPT-5.4 [xhigh] | 52% | $5.65 | 71k | 70 |
| 19 | Gemini 3.6 Flash [high] | 47% | $4.42 | 96k | 117 |
| 20 | GLM 5.2 [max] | 44% | $3.92 | 78k | 129 |
| 21 | Gemini 3.5 Flash [high] | 36% | $3.45 | 76k | 105 |
| 22 | Kimi K2.7 Code | 31% | $2.82 | 59k | 149 |
| 23 | Claude Sonnet 4.6 [high] | 30% | $5.52 | 76k | 134 |
| 24 | Gemini 3.1 Pro [high] | 12% | $2.14 | 28k | 76 |
DeepSWE “Best” picks the highest published reasoning effort per model (not the highest pass rate). Small gaps may not be statistically meaningful — confirm on deepswe.datacurve.ai.
Head to head
Direct comparisons
Pairs a buyer would realistically weigh against each other: same tier, or a step up and down within one provider.
GPT-5.6 Sol vs GPT-5.6 Terra
GPT-5.6 Luna vs GPT-5.6 Sol
Claude Opus 5 vs GPT-5.6 Sol
Claude Fable 5 vs GPT-5.6 Sol
Gemini 3.1 Pro vs GPT-5.6 Sol
GPT-5.6 Sol vs Kimi K3
GLM 5.2 vs GPT-5.6 Sol
GLM 5.3 vs GPT-5.6 Sol
GPT-5.6 Sol vs Grok 4.6
GPT-5.6 Sol vs Qwen3.8-Max
GPT-5.6 Sol vs Qwen3.8-2.4T-A95B
GPT-5.6 Sol vs Muse Spark 1.2
GPT-5.6 Luna vs GPT-5.6 Terra
Claude Sonnet 5 vs GPT-5.6 Terra
GPT-5.6 Terra vs Grok 4.5
DeepSeek V4-Pro vs GPT-5.6 Terra
GPT-5.6 Terra vs Qwen3.8-27B
GPT-5.6 Terra vs Muse Glimmer 30B
Gemini 3.6 Flash vs GPT-5.6 Luna
Gemini 3.7 Flash vs GPT-5.6 Luna
Gemini 3.5 Flash-Lite vs GPT-5.6 Luna
GPT-5.6 Luna vs Kimi K2.6
Claude Haiku 4.5 vs GPT-5.6 Luna
DeepSeek V4-Flash vs GPT-5.6 Luna
GPT-5.6 Luna vs Ling 3.0 Flash
Claude Fable 5 vs Claude Opus 5
Claude Opus 5 vs Claude Sonnet 5
Claude Opus 5 vs Gemini 3.1 Pro
Claude Opus 5 vs Kimi K3
Claude Opus 5 vs GLM 5.2
Claude Opus 5 vs GLM 5.3
Claude Opus 5 vs Grok 4.6
Claude Haiku 4.5 vs Claude Opus 5
Claude Opus 5 vs Qwen3.8-Max
Claude Opus 5 vs Qwen3.8-2.4T-A95B
Claude Opus 5 vs Muse Spark 1.2
Claude Fable 5 vs Claude Sonnet 5
Claude Fable 5 vs Gemini 3.1 Pro
Claude Fable 5 vs Kimi K3
Claude Fable 5 vs GLM 5.2
Claude Fable 5 vs GLM 5.3
Claude Fable 5 vs Grok 4.6
Claude Fable 5 vs Claude Haiku 4.5
Claude Fable 5 vs Qwen3.8-Max
Claude Fable 5 vs Qwen3.8-2.4T-A95B
Claude Fable 5 vs Muse Spark 1.2
Claude Sonnet 5 vs Grok 4.5
Claude Haiku 4.5 vs Claude Sonnet 5
Claude Sonnet 5 vs DeepSeek V4-Pro
Claude Sonnet 5 vs Qwen3.8-27B
Claude Sonnet 5 vs Muse Glimmer 30B
Gemini 3.1 Pro vs Gemini 3.6 Flash
Gemini 3.1 Pro vs Gemini 3.7 Flash
Gemini 3.1 Pro vs Gemini 3.5 Flash-Lite
Gemini 3.1 Pro vs Kimi K3
Gemini 3.1 Pro vs GLM 5.2
Gemini 3.1 Pro vs GLM 5.3
Gemini 3.1 Pro vs Grok 4.6
Gemini 3.1 Pro vs Qwen3.8-Max
Gemini 3.1 Pro vs Qwen3.8-2.4T-A95B
Gemini 3.1 Pro vs Muse Spark 1.2
Gemini 3.6 Flash vs Gemini 3.7 Flash
Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash
Gemini 3.6 Flash vs Kimi K2.6
Claude Haiku 4.5 vs Gemini 3.6 Flash
DeepSeek V4-Flash vs Gemini 3.6 Flash
Gemini 3.6 Flash vs Ling 3.0 Flash
Gemini 3.5 Flash-Lite vs Gemini 3.7 Flash
Gemini 3.7 Flash vs Kimi K2.6
Claude Haiku 4.5 vs Gemini 3.7 Flash
DeepSeek V4-Flash vs Gemini 3.7 Flash
Gemini 3.7 Flash vs Ling 3.0 Flash
Gemini 3.5 Flash-Lite vs Kimi K2.6
Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite
DeepSeek V4-Flash vs Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite vs Ling 3.0 Flash
Kimi K2.6 vs Kimi K3
GLM 5.2 vs Kimi K3
GLM 5.3 vs Kimi K3
Grok 4.6 vs Kimi K3
Kimi K3 vs Qwen3.8-Max
Kimi K3 vs Qwen3.8-2.4T-A95B
Kimi K3 vs Muse Spark 1.2
Claude Haiku 4.5 vs Kimi K2.6
DeepSeek V4-Flash vs Kimi K2.6
Kimi K2.6 vs Ling 3.0 Flash
GLM 5.2 vs GLM 5.3
GLM 5.2 vs Grok 4.6
GLM 5.2 vs Qwen3.8-Max
GLM 5.2 vs Qwen3.8-2.4T-A95B
GLM 5.2 vs Muse Spark 1.2
GLM 5.3 vs Grok 4.6
GLM 5.3 vs Qwen3.8-Max
GLM 5.3 vs Qwen3.8-2.4T-A95B
GLM 5.3 vs Muse Spark 1.2
Grok 4.5 vs Grok 4.6
DeepSeek V4-Pro vs Grok 4.5
Grok 4.5 vs Qwen3.8-27B
Grok 4.5 vs Muse Glimmer 30B
Grok 4.6 vs Qwen3.8-Max
Grok 4.6 vs Qwen3.8-2.4T-A95B
Grok 4.6 vs Muse Spark 1.2
Claude Haiku 4.5 vs DeepSeek V4-Flash
Claude Haiku 4.5 vs Ling 3.0 Flash
DeepSeek V4-Flash vs DeepSeek V4-Pro
DeepSeek V4-Flash vs Ling 3.0 Flash
DeepSeek V4-Pro vs Qwen3.8-27B
DeepSeek V4-Pro vs Muse Glimmer 30B
Qwen3.8-2.4T-A95B vs Qwen3.8-Max
Qwen3.8-27B vs Qwen3.8-Max
Muse Spark 1.2 vs Qwen3.8-Max
Qwen3.8-2.4T-A95B vs Qwen3.8-27B
Muse Spark 1.2 vs Qwen3.8-2.4T-A95B
Muse Glimmer 30B vs Qwen3.8-27B
Muse Glimmer 30B vs Muse Spark 1.2
Claude Opus 5 vs Gemini 3.6 Flash
Claude Opus 5 vs Gemini 3.7 Flash
Claude Opus 5 vs DeepSeek V4-Flash
Claude Opus 5 vs GPT-5.6 Luna
DeepSeek V4-Flash vs GPT-5.6 Sol
Gemini 3.7 Flash vs GPT-5.6 Sol
GLM 5.2 vs Grok 4.5
GLM 5.3 vs Grok 4.5
How these comparisons are built
Every page renders from the same checked catalog dataset rather than a written template.
A comparison only exists where the underlying numbers differ. Figures we could not confirm from the provider are shown as Not verified rather than estimated, because a wrong price is more damaging than a missing one.
Benchmark results carry the source and the month they were measured. Scores in this category move with every release, and a number without a date is not much use.
Picking a tier
The tier usually matters more than the brand.
Frontier tiers earn their price on hard reasoning and long agentic tasks. For classification, extraction and summarization — most production volume — a budget tier is usually indistinguishable in output and several times cheaper.
Output tokens cost three to five times input across almost every provider here, so generation length drives the bill far more than prompt size.