Skip to main content
← Benchmark sources

Model providers · benchmark source

Provider catalogs

System limits from provider documentation — context capacity is not a quality benchmark.

Open Model providers1 board on this publisher

System capacity

Context window

The amount of input and output context a model can process in one interaction under the provider's rules.

Why it matters

It sets a hard ceiling for documents, repositories, and conversation state.

Use it when

You are sizing a real workload and will test long-context quality, truncation, and cost.

The limitation

More tokens do not guarantee better retrieval, attention, or reasoning across the whole window.

Leaderboard snapshot

Context window results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

Local leader

GPT-5.6 Sol

1.05M

Rows shown

25

Context tokens

Snapshot date

2026-08

14 days old · not a live API feed

Score profile

Context tokens by published row

More capacity is not automatically better

GPT-5.6 Sol
1.05M
GPT-5.6 Terra
1.05M
GPT-5.6 Luna
1.05M
Gemini 3.1 Pro
1.05M
Gemini 3.6 Flash
1.05M
Gemini 3.7 Flash
1.05M
Gemini 3.5 Flash-Lite
1.05M
Kimi K3
1.05M
Muse Spark 1.2
1.05M
Claude Opus 5
1M
Claude Fable 5
1M
Claude Sonnet 5
1M
GLM 5.2
1M
GLM 5.3
1M
DeepSeek V4-Flash
1M
DeepSeek V4-Pro
1M
Qwen3.8-Max
1M
Grok 4.5
500K
Grok 4.6
500K
Kimi K2.6
262K
Ling 3.0 Flash
262K
Qwen3.8-2.4T-A95B
262K
Qwen3.8-27B
262K
Claude Haiku 4.5
200K
Muse Glimmer 30B
131K

Provider-catalogued context capacities for the models in this site’s dataset. Capacity is a system limit, not a quality benchmark; the linked OpenAI documentation is an example receipt, not a source for every provider row.

Open Provider documentation example (OpenAI)

Evidence in this catalog

Where Context window fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

No directly comparable model score is verified here yet.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on Provider documentation example (OpenAI)