swebench.com · benchmark source
SWE-bench
Repository-level GitHub issue resolution benchmarks. Verified is a human-filtered historical subset — treat frontier progress claims carefully.
Software engineering
SWE-bench Verified
A historical benchmark of real GitHub issues where an agent must modify a repository so tests and issue requirements are satisfied.
Why it matters
It gives coding agents a shared task language closer to real maintenance than a short code-generation prompt.
Use it when
You are reproducing historical repository-level results and can document the same harness assumptions; use a less contaminated evaluation for current progress.
The limitation
OpenAI reports contamination and grading flaws in SWE-bench Verified, so it should not be treated as current frontier-progress evidence. Harness, scaffold, test selection, patch filtering, and model effort also materially change the score.
Leaderboard snapshot
SWE-bench Verified results in a readable view
Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.
Local leader
Claude Opus 5
96%
Rows shown
6
Resolved
Snapshot date
2026-08-11
4 days old · not a live API feed
Score profile
Resolved by published row
Higher is better in this view
Local leader in this group: Claude Opus 5 (96%)
Score profile
Resolved by published row
Higher is better in this view
Local leader in this group: Kimi K2.6 (80.2%)
| Rank | Model or system | Resolved | Other details | Note |
|---|---|---|---|---|
| #1 | 96% | Source: Provider / BenchLM Aug 2026Harness: Provider scaffold (not the official SWE-bench agent)Trials: Avg of 5 (system card) | Saturated historical board; OpenAI cautions Verified for frontier progress claims | |
| #2 | 95% | Source: Provider / BenchLM Aug 2026Harness: Provider scaffold (not the official SWE-bench agent) | — | |
| #3 | 85.2% | Source: Provider / BenchLM Aug 2026Harness: Provider scaffold (not the official SWE-bench agent) | — | |
| #4 | 80.2% | Source: BenchLM aggregator Aug 2026Harness: Aggregator-reported; scaffold not independently reproduced here | — | |
| #5 | 79% | Source: Provider / BenchLM Aug 2026Harness: Provider scaffold (not the official SWE-bench agent) | — | |
| #6 | 73.3% | Source: BenchLM aggregator Aug 2026Harness: Aggregator-reported; scaffold not independently reproduced here | — |
Orienting shortlist of catalogued models with published SWE-bench Verified resolved rates (retrieved 2026-08-11 via public/provider aggregation). Scaffolding and trial averaging differ by row; OpenAI also cautions Verified for current frontier-progress claims. Open the official board for the complete live table.
Open SWE-bench Verified leaderboardEvidence in this catalog
Where SWE-bench Verified fits
Closest verified examples we currently carry. A missing score is not a zero.
Model examples
- Claude Opus 596% resolved
Provider-reported Verified result (avg of 5 trials); scaffolding still matters.
- Claude Fable 595% resolved
Catalogued Verified shortlist row; check the live board for harness variants.
- Claude Sonnet 585.2% resolved
Catalogued Verified shortlist row for the mid-tier Claude option.
Harness examples
No directly comparable harness score is verified here yet.