Skip to main content
← Benchmark sources

livebench.ai · benchmark source

LiveBench

A contamination-aware suite that refreshes questions over time. Category and subtask views live on the source site.

Open livebench.ai1 board on this publisher

Fresh evaluation

LiveBench overall

A benchmark family designed to refresh questions and reduce contamination from training data.

Why it matters

Fresh tasks can provide a less stale view of reasoning and language performance.

Use it when

You want a current external signal alongside static benchmark suites.

The limitation

It still measures selected tasks and cannot predict your full workflow or tool integration.

Leaderboard snapshot

LiveBench overall results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

Full LiveBench 2026-06-25 overall table mirrored from livebench.ai (43 published overall rows), including cost per successful task.

Local leader

Claude Fable 5 Max Effort

83%

Rows shown

43

Overall score

Snapshot date

2026-06-25

51 days old · not a live API feed

Open the live source before trusting rank order — especially when this snapshot predates recent model releases.

Score profile

Overall score by published row

Higher is better in this view

Claude Fable 5 Max Effort
83%
GPT-5.6 Sol Max Effort
81%
GPT-5.5 Thinking xHigh Effort
80.2%
Claude 5 Opus Thinking Max Effort
80.1%
Smaug-Agentic
79.5%
Kimi K3
79.2%
Gemini 3.7 Flash High
78.8%
Qwen 3.8 Max
78.5%
GPT-5.4 Thinking xHigh Effort
78%
Grok 4.6
78%
Muse Spark 1.2 xHigh Effort
78%
GPT-5.6 Terra Max Effort
77.9%
DeepSeek V4 Pro 0813
77.4%
Gemini 3.1 Pro Preview High
77%
Claude 4.7 Opus Thinking xHigh Effort
76.5%
Claude 4.8 Opus Thinking Max Effort
76.2%
Claude Sonnet 5 xHigh Effort
76%
Grok 4.5
75.8%
Muse Spark 1.1 xHigh Effort
75.3%
Gemini 3.5 Flash High
74.6%
GPT-5.2 High
74.6%
Claude 4.6 Opus Thinking High Effort
74.5%
DeepSeek V4 Flash 0731
74.2%
GPT-5.2 Codex
74%
Gemini 3.6 Flash High
73.6%
GPT-5.6 Luna Max Effort
73.6%
GLM-5.2
73.2%
Qwen 3.7 Max
73.1%
Claude 4.6 Sonnet Thinking Medium Effort
73%
Claude 4.5 Opus Thinking High Effort
72.6%
Inkling xHigh Effort
71.9%
DeepSeek V4 Pro
71.6%
Kimi K2.6 Thinking
70.5%
GPT-5.4 Nano xHigh
69.6%
Qwen 3.6 Plus
68.9%
Kimi K2.7 Code
68.4%
Grok Build 0.1
67.8%
Minimax M3
67.3%
GPT-5.4 Mini xHigh
66.4%
DeepSeek V4 Flash
65.5%
Qwen 3.6 27B
64%
Gemini 3.5 Flash-Lite High
63.9%
Grok 4.3
62.3%

Full LiveBench 2026-06-25 overall release mirrored from livebench.ai (43 published overall rows with category columns available on the source). LiveBench rotates questions; open the source for category/subtask detail and later releases.

Open LiveBench leaderboard

Evidence in this catalog

Where LiveBench overall fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

No directly comparable model score is verified here yet.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on LiveBench