Skip to main content

Compare models

Price, context, and benchmarks side by side

Context window, price per million tokens, reasoning levels, and benchmark results with the date they were checked. Unconfirmed fields are marked; provider pricing and capability details can change.

This is the full comparison table for model APIs (GPT, Claude, Gemini, and so on). For launches in the last 28 days, open New models. For Cursor, Claude Code, Muse Code, Lovable and similar products, see Agentic harnesses. Right now that includes 13 recent launches.

Run the model picker4 questions → frontier, balanced, or budget shortlist

Task guides

Choose by the work you need done

Short guides for coding, writing, research, RAG, voice, and more — then compare the models that actually fit.

AI model tier, context, output, pricing, and weight comparison
ModelTierContextMax outputIn / 1MOut / 1MWeights
GPT-5.6 SolOpenAIFrontier1.05M128K$5$30Closed
GPT-5.6 TerraOpenAIBalanced1.05M128K$2$12Closed
GPT-5.6 LunaOpenAIBudget1.05M128K$0.20$1.20Closed
Claude Opus 5AnthropicFrontier1M128K$5$25Closed
Claude Fable 5AnthropicFrontier1M128K$10$50Closed
Claude Sonnet 5AnthropicBalanced1M128K$2$10Closed
Gemini 3.1 ProGoogleFrontier1.05M66K$2$12Closed
Gemini 3.6 FlashGoogleBudget1.05M66K$0.75$3.75Closed
Gemini 3.7 FlashGoogleBudget1.05M66K$0.75$3.75Closed
Gemini 3.5 Flash-LiteGoogleBudget1.05M66K$0.30$2.50Closed
Kimi K3Moonshot AIFrontier1.05M1M$3$15Open
Kimi K2.6Moonshot AIBudget262KNot verified$0.95$4Open
GLM 5.2Z.aiFrontier1M128K$1.40$4.40Open
GLM 5.3Z.aiFrontier1M128KNot verifiedNot verifiedClosed
Grok 4.5SpaceXAIBalanced500KNot verified$2$6Closed
Grok 4.6SpaceXAIFrontier500KNot verified$2$6Closed
Claude Haiku 4.5AnthropicBudget200K64K$1$5Closed
DeepSeek V4-FlashDeepSeekBudget1M384K$0.14$0.28Open
DeepSeek V4-ProDeepSeekBalanced1M384K$0.435$0.87Open
Ling 3.0 FlashInclusionAIBudget262KNot verified$0.075$0.22Open
Qwen3.8-MaxQwenFrontier1M131K$2$6Closed
Qwen3.8-2.4T-A95BQwenFrontier262K131KNot verifiedNot verifiedOpen
Qwen3.8-27BQwenBalanced262K131KNot verifiedNot verifiedOpen
Muse Spark 1.2MetaFrontier1.05MNot verified$1.25$4.25Closed
Muse Glimmer 30BMetaBalanced131KNot verified$0$0Open

Compare two models

Open comparison

Prices are USD per million tokens at standard rates, excluding batch and caching discounts. Last verified August 14, 2026. Model pricing and capability in this category change frequently — check the provider before committing spend.

Benchmark

DeepSWE 1.1 — coding agent performance and cost

Score alone hides the decision. Cost per completed task sits alongside it, because a model a few points behind at a fraction of the price is usually the better build.

Local leader

Claude Opus 5 [max]

74%

Rows shown

24

Highest published reasoning effort per model (not best Pass@1)

Snapshot date

2026-08-13

Mirrored from deepswe.datacurve.ai

DeepSWE 1.1 pass@10%16%32%48%64%80%$0$4.50$9.00$13.50$18.00$22.50$27.00Avg cost per taskClaude Opus 5 [max]: 74% · $11.84 · 118k tokens · 99 stepsClaude Opus 5GPT-5.6 Sol [max]: 73% · $8.39 · 60k tokens · 61 stepsGPT-5.6 SolClaude Fable 5 [max]: 70% · $21.63 · 119k tokens · 88 stepsClaude Fable 5GPT-5.6 Terra [max]: 70% · $4.95 · 72k tokens · 76 stepsGPT-5.6 TerraKimi K3 [max]: 69% · $4.65 · 82k tokens · 98 stepsKimi K3GPT-5.6 Luna [max]: 67% · $3.03 · 73k tokens · 102 stepsGPT-5.6 LunaGPT-5.5 [xhigh]: 67% · $7.23 · 46k tokens · 82 stepsGPT-5.5Grok 4.6 [xhigh]: 67% · $5.50 · 71k tokens · 87 stepsGrok 4.6Gemini 3.7 Flash [high]: 65% · $2.18 · 107k tokens · 125 stepsGemini 3.7 FlashDeepSeek V4-Pro [max]: 63% · $0.24 · 106k tokens · 155 stepsDeepSeek V4-ProClaude Opus 4.8 [max]: 59% · $13.22 · 135k tokens · 120 stepsClaude Opus 4.8Qwen3.8-Max [xhigh]: 58% · $3.73 · 95k tokens · 111 stepsQwen3.8-MaxMuse Spark 1.2 [xhigh]: 55% · $3.70 · 99k tokens · 101 stepsMuse Spark 1.2Claude Sonnet 5 [max]: 54% · $26.40 · 214k tokens · 268 stepsClaude Sonnet 5Grok 4.5 [high]: 54% · $2.42 · 36k tokens · 61 stepsGrok 4.5DeepSeek V4-Flash [max]: 53% · $0.10 · 108k tokens · 153 stepsDeepSeek V4-FlashMuse Spark 1.1 [xhigh]: 53% · $2.36 · 74k tokens · 96 stepsMuse Spark 1.1GPT-5.4 [xhigh]: 52% · $5.65 · 71k tokens · 70 stepsGPT-5.4Gemini 3.6 Flash [high]: 47% · $4.42 · 96k tokens · 117 stepsGemini 3.6 FlashGLM 5.2 [max]: 44% · $3.92 · 78k tokens · 129 stepsGLM 5.2Gemini 3.5 Flash [high]: 36% · $3.45 · 76k tokens · 105 stepsGemini 3.5 FlashKimi K2.7 Code: 31% · $2.82 · 59k tokens · 149 stepsKimi K2.7 CodeClaude Sonnet 4.6 [high]: 30% · $5.52 · 76k tokens · 134 stepsClaude Sonnet 4.6Gemini 3.1 Pro [high]: 12% · $2.14 · 28k tokens · 76 stepsGemini 3.1 Pro

Better is toward the top-right (higher pass rate, lower cost). X-axis is reversed to match DeepSWE’s public chart. v1.1 uses average cost / tokens / steps; v1 uses published medians.

DeepSWE 1.1 leaderboard with pass rate, cost, tokens, and steps per task
#ModelPass@1Cost / taskTokens / taskSteps / task
1Claude Opus 5 [max]74%$11.84118k99
2GPT-5.6 Sol [max]73%$8.3960k61
3Claude Fable 5 [max]70%$21.63119k88
4GPT-5.6 Terra [max]70%$4.9572k76
5Kimi K3 [max]69%$4.6582k98
6GPT-5.6 Luna [max]67%$3.0373k102
7GPT-5.5 [xhigh]67%$7.2346k82
8Grok 4.6 [xhigh]67%$5.5071k87
9Gemini 3.7 Flash [high]65%$2.18107k125
10DeepSeek V4-Pro [max]63%$0.24106k155
11Claude Opus 4.8 [max]59%$13.22135k120
12Qwen3.8-Max [xhigh]58%$3.7395k111
13Muse Spark 1.2 [xhigh]55%$3.7099k101
14Claude Sonnet 5 [max]54%$26.40214k268
15Grok 4.5 [high]54%$2.4236k61
16DeepSeek V4-Flash [max]53%$0.10108k153
17Muse Spark 1.1 [xhigh]53%$2.3674k96
18GPT-5.4 [xhigh]52%$5.6571k70
19Gemini 3.6 Flash [high]47%$4.4296k117
20GLM 5.2 [max]44%$3.9278k129
21Gemini 3.5 Flash [high]36%$3.4576k105
22Kimi K2.7 Code31%$2.8259k149
23Claude Sonnet 4.6 [high]30%$5.5276k134
24Gemini 3.1 Pro [high]12%$2.1428k76

DeepSWE “Best” picks the highest published reasoning effort per model (not the highest pass rate). Small gaps may not be statistically meaningful — confirm on deepswe.datacurve.ai.

Head to head

Direct comparisons

Pairs a buyer would realistically weigh against each other: same tier, or a step up and down within one provider.

vs

GPT-5.6 Sol vs GPT-5.6 Terra

vs

GPT-5.6 Luna vs GPT-5.6 Sol

vs

Claude Opus 5 vs GPT-5.6 Sol

vs

Claude Fable 5 vs GPT-5.6 Sol

vs

Gemini 3.1 Pro vs GPT-5.6 Sol

vs

GPT-5.6 Sol vs Kimi K3

vs

GLM 5.2 vs GPT-5.6 Sol

vs

GLM 5.3 vs GPT-5.6 Sol

vs

GPT-5.6 Sol vs Grok 4.6

vs

GPT-5.6 Sol vs Qwen3.8-Max

vs

GPT-5.6 Sol vs Qwen3.8-2.4T-A95B

vs

GPT-5.6 Sol vs Muse Spark 1.2

vs

GPT-5.6 Luna vs GPT-5.6 Terra

vs

Claude Sonnet 5 vs GPT-5.6 Terra

vs

GPT-5.6 Terra vs Grok 4.5

vs

DeepSeek V4-Pro vs GPT-5.6 Terra

vs

GPT-5.6 Terra vs Qwen3.8-27B

vs

GPT-5.6 Terra vs Muse Glimmer 30B

vs

Gemini 3.6 Flash vs GPT-5.6 Luna

vs

Gemini 3.7 Flash vs GPT-5.6 Luna

vs

Gemini 3.5 Flash-Lite vs GPT-5.6 Luna

vs

GPT-5.6 Luna vs Kimi K2.6

vs

Claude Haiku 4.5 vs GPT-5.6 Luna

vs

DeepSeek V4-Flash vs GPT-5.6 Luna

vs

GPT-5.6 Luna vs Ling 3.0 Flash

vs

Claude Fable 5 vs Claude Opus 5

vs

Claude Opus 5 vs Claude Sonnet 5

vs

Claude Opus 5 vs Gemini 3.1 Pro

vs

Claude Opus 5 vs Kimi K3

vs

Claude Opus 5 vs GLM 5.2

vs

Claude Opus 5 vs GLM 5.3

vs

Claude Opus 5 vs Grok 4.6

vs

Claude Haiku 4.5 vs Claude Opus 5

vs

Claude Opus 5 vs Qwen3.8-Max

vs

Claude Opus 5 vs Qwen3.8-2.4T-A95B

vs

Claude Opus 5 vs Muse Spark 1.2

vs

Claude Fable 5 vs Claude Sonnet 5

vs

Claude Fable 5 vs Gemini 3.1 Pro

vs

Claude Fable 5 vs Kimi K3

vs

Claude Fable 5 vs GLM 5.2

vs

Claude Fable 5 vs GLM 5.3

vs

Claude Fable 5 vs Grok 4.6

vs

Claude Fable 5 vs Claude Haiku 4.5

vs

Claude Fable 5 vs Qwen3.8-Max

vs

Claude Fable 5 vs Qwen3.8-2.4T-A95B

vs

Claude Fable 5 vs Muse Spark 1.2

vs

Claude Sonnet 5 vs Grok 4.5

vs

Claude Haiku 4.5 vs Claude Sonnet 5

vs

Claude Sonnet 5 vs DeepSeek V4-Pro

vs

Claude Sonnet 5 vs Qwen3.8-27B

vs

Claude Sonnet 5 vs Muse Glimmer 30B

vs

Gemini 3.1 Pro vs Gemini 3.6 Flash

vs

Gemini 3.1 Pro vs Gemini 3.7 Flash

vs

Gemini 3.1 Pro vs Gemini 3.5 Flash-Lite

vs

Gemini 3.1 Pro vs Kimi K3

vs

Gemini 3.1 Pro vs GLM 5.2

vs

Gemini 3.1 Pro vs GLM 5.3

vs

Gemini 3.1 Pro vs Grok 4.6

vs

Gemini 3.1 Pro vs Qwen3.8-Max

vs

Gemini 3.1 Pro vs Qwen3.8-2.4T-A95B

vs

Gemini 3.1 Pro vs Muse Spark 1.2

vs

Gemini 3.6 Flash vs Gemini 3.7 Flash

vs

Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash

vs

Gemini 3.6 Flash vs Kimi K2.6

vs

Claude Haiku 4.5 vs Gemini 3.6 Flash

vs

DeepSeek V4-Flash vs Gemini 3.6 Flash

vs

Gemini 3.6 Flash vs Ling 3.0 Flash

vs

Gemini 3.5 Flash-Lite vs Gemini 3.7 Flash

vs

Gemini 3.7 Flash vs Kimi K2.6

vs

Claude Haiku 4.5 vs Gemini 3.7 Flash

vs

DeepSeek V4-Flash vs Gemini 3.7 Flash

vs

Gemini 3.7 Flash vs Ling 3.0 Flash

vs

Gemini 3.5 Flash-Lite vs Kimi K2.6

vs

Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

vs

DeepSeek V4-Flash vs Gemini 3.5 Flash-Lite

vs

Gemini 3.5 Flash-Lite vs Ling 3.0 Flash

vs

Kimi K2.6 vs Kimi K3

vs

GLM 5.2 vs Kimi K3

vs

GLM 5.3 vs Kimi K3

vs

Grok 4.6 vs Kimi K3

vs

Kimi K3 vs Qwen3.8-Max

vs

Kimi K3 vs Qwen3.8-2.4T-A95B

vs

Kimi K3 vs Muse Spark 1.2

vs

Claude Haiku 4.5 vs Kimi K2.6

vs

DeepSeek V4-Flash vs Kimi K2.6

vs

Kimi K2.6 vs Ling 3.0 Flash

vs

GLM 5.2 vs GLM 5.3

vs

GLM 5.2 vs Grok 4.6

vs

GLM 5.2 vs Qwen3.8-Max

vs

GLM 5.2 vs Qwen3.8-2.4T-A95B

vs

GLM 5.2 vs Muse Spark 1.2

vs

GLM 5.3 vs Grok 4.6

vs

GLM 5.3 vs Qwen3.8-Max

vs

GLM 5.3 vs Qwen3.8-2.4T-A95B

vs

GLM 5.3 vs Muse Spark 1.2

vs

Grok 4.5 vs Grok 4.6

vs

DeepSeek V4-Pro vs Grok 4.5

vs

Grok 4.5 vs Qwen3.8-27B

vs

Grok 4.5 vs Muse Glimmer 30B

vs

Grok 4.6 vs Qwen3.8-Max

vs

Grok 4.6 vs Qwen3.8-2.4T-A95B

vs

Grok 4.6 vs Muse Spark 1.2

vs

Claude Haiku 4.5 vs DeepSeek V4-Flash

vs

Claude Haiku 4.5 vs Ling 3.0 Flash

vs

DeepSeek V4-Flash vs DeepSeek V4-Pro

vs

DeepSeek V4-Flash vs Ling 3.0 Flash

vs

DeepSeek V4-Pro vs Qwen3.8-27B

vs

DeepSeek V4-Pro vs Muse Glimmer 30B

vs

Qwen3.8-2.4T-A95B vs Qwen3.8-Max

vs

Qwen3.8-27B vs Qwen3.8-Max

vs

Muse Spark 1.2 vs Qwen3.8-Max

vs

Qwen3.8-2.4T-A95B vs Qwen3.8-27B

vs

Muse Spark 1.2 vs Qwen3.8-2.4T-A95B

vs

Muse Glimmer 30B vs Qwen3.8-27B

vs

Muse Glimmer 30B vs Muse Spark 1.2

vs

Claude Opus 5 vs Gemini 3.6 Flash

vs

Claude Opus 5 vs Gemini 3.7 Flash

vs

Claude Opus 5 vs DeepSeek V4-Flash

vs

Claude Opus 5 vs GPT-5.6 Luna

vs

DeepSeek V4-Flash vs GPT-5.6 Sol

vs

Gemini 3.7 Flash vs GPT-5.6 Sol

vs

GLM 5.2 vs Grok 4.5

vs

GLM 5.3 vs Grok 4.5

How these comparisons are built

Every page renders from the same checked catalog dataset rather than a written template.

A comparison only exists where the underlying numbers differ. Figures we could not confirm from the provider are shown as Not verified rather than estimated, because a wrong price is more damaging than a missing one.

Benchmark results carry the source and the month they were measured. Scores in this category move with every release, and a number without a date is not much use.

Picking a tier

The tier usually matters more than the brand.

FrontierBalancedBudget

Frontier tiers earn their price on hard reasoning and long agentic tasks. For classification, extraction and summarization — most production volume — a budget tier is usually indistinguishable in output and several times cheaper.

Output tokens cost three to five times input across almost every provider here, so generation length drives the bill far more than prompt size.