livebench.ai · benchmark source
LiveBench
A contamination-aware suite that refreshes questions over time. Category and subtask views live on the source site.
Fresh evaluation
LiveBench overall
A benchmark family designed to refresh questions and reduce contamination from training data.
Why it matters
Fresh tasks can provide a less stale view of reasoning and language performance.
Use it when
You want a current external signal alongside static benchmark suites.
The limitation
It still measures selected tasks and cannot predict your full workflow or tool integration.
Leaderboard snapshot
LiveBench overall results in a readable view
Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.
Local leader
Claude Fable 5 Max Effort
83%
Rows shown
43
Overall score
Snapshot date
2026-06-25
51 days old · not a live API feed
Open the live source before trusting rank order — especially when this snapshot predates recent model releases.
Score profile
Overall score by published row
Higher is better in this view
Source detail
Cost / task from the published rows
Sorted by this field · not the primary ranking
| Rank | Model or system | Overall score | Cost / task | Other details | Note |
|---|---|---|---|---|---|
| #1 | 83% | $1.439 | — | — | |
| #2 | 81% | $0.515 | — | — | |
| #3 | 80.2% | $0.435 | — | — | |
| #4 | 80.1% | $0.699 | — | — | |
| #5 | 79.5% | $0.329 | — | — | |
| #6 | 79.2% | $0.348 | — | — | |
| #7 | 78.8% | $0.157 | — | — | |
| #8 | 78.5% | $0.275 | — | — | |
| #9 | 78% | $0.387 | — | — | |
| #9 | 78% | $0.207 | — | — | |
| #9 | 78% | $0.375 | — | — | |
| #12 | 77.9% | $0.352 | — | — | |
| #13 | 77.4% | $0.044 | — | — | |
| #14 | 77% | $0.286 | — | — | |
| #15 | 76.5% | $0.528 | — | — | |
| #16 | 76.2% | $0.983 | — | — | |
| #17 | 76% | $0.505 | — | — | |
| #18 | 75.8% | $0.131 | — | — | |
| #19 | 75.3% | $0.198 | — | — | |
| #20 | 74.6% | $0.249 | — | — | |
| #20 | 74.6% | $0.234 | — | — | |
| #22 | 74.5% | $0.404 | — | — | |
| #23 | 74.2% | $0.060 | — | — | |
| #24 | 74% | $0.187 | — | — | |
| #25 | 73.6% | $0.235 | — | — | |
| #25 | 73.6% | $0.169 | — | — | |
| #27 | 73.2% | $0.225 | — | — | |
| #28 | 73.1% | $0.182 | — | — | |
| #29 | 73% | $0.306 | — | — | |
| #30 | 72.6% | $0.610 | — | — | |
| #31 | 71.9% | $0.310 | — | — | |
| #32 | 71.6% | $0.050 | — | — | |
| #33 | 70.5% | $0.169 | — | — | |
| #34 | 69.6% | $0.091 | — | — | |
| #35 | 68.9% | $0.227 | — | — | |
| #36 | 68.4% | $0.100 | — | — | |
| #37 | 67.8% | $0.024 | — | — | |
| #38 | 67.3% | $0.060 | — | — | |
| #39 | 66.4% | $0.334 | — | — | |
| #40 | 65.5% | $0.016 | — | — | |
| #41 | 64% | $0.202 | — | — | |
| #42 | 63.9% | $0.069 | — | — | |
| #43 | 62.3% | $0.061 | — | — |
Full LiveBench 2026-06-25 overall release mirrored from livebench.ai (43 published overall rows with category columns available on the source). LiveBench rotates questions; open the source for category/subtask detail and later releases.
Open LiveBench leaderboardEvidence in this catalog
Where LiveBench overall fits
Closest verified examples we currently carry. A missing score is not a zero.
Model examples
No directly comparable model score is verified here yet.
Harness examples
No directly comparable harness score is verified here yet.