Skip to main content
← Benchmark sources

swebench.com · benchmark source

SWE-bench

Repository-level GitHub issue resolution benchmarks. Verified is a human-filtered historical subset — treat frontier progress claims carefully.

Open swebench.com1 board on this publisher

Software engineering

SWE-bench Verified

A historical benchmark of real GitHub issues where an agent must modify a repository so tests and issue requirements are satisfied.

Why it matters

It gives coding agents a shared task language closer to real maintenance than a short code-generation prompt.

Use it when

You are reproducing historical repository-level results and can document the same harness assumptions; use a less contaminated evaluation for current progress.

The limitation

OpenAI reports contamination and grading flaws in SWE-bench Verified, so it should not be treated as current frontier-progress evidence. Harness, scaffold, test selection, patch filtering, and model effort also materially change the score.

Leaderboard snapshot

SWE-bench Verified results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

SWE-bench Verified is saturated for frontier models. OpenAI cautions against treating it as evidence of current frontier progress — scaffolding, trial averaging, and effort settings also move the score materially.
Official swebench.com remains JS-rendered; this page keeps a verified catalog shortlist until a stable full-table scrape is available.
Rows come from provider cards and aggregators, not one official harness. Groups below are provenance families — do not read the list as a single comparable ranking.

Local leader

Claude Opus 5

96%

Rows shown

6

Resolved

Snapshot date

2026-08-11

4 days old · not a live API feed

Score profile

Resolved by published row

Higher is better in this view

Claude Opus 5
96%
Claude Fable 5
95%
Claude Sonnet 5
85.2%
DeepSeek V4-Flash [max]
79%

Local leader in this group: Claude Opus 5 (96%)

Score profile

Resolved by published row

Higher is better in this view

Kimi K2.6
80.2%
Claude Haiku 4.5
73.3%

Local leader in this group: Kimi K2.6 (80.2%)

Orienting shortlist of catalogued models with published SWE-bench Verified resolved rates (retrieved 2026-08-11 via public/provider aggregation). Scaffolding and trial averaging differ by row; OpenAI also cautions Verified for current frontier-progress claims. Open the official board for the complete live table.

Open SWE-bench Verified leaderboard

Evidence in this catalog

Where SWE-bench Verified fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

  • Claude Opus 596% resolved

    Provider-reported Verified result (avg of 5 trials); scaffolding still matters.

  • Claude Fable 595% resolved

    Catalogued Verified shortlist row; check the live board for harness variants.

  • Claude Sonnet 585.2% resolved

    Catalogued Verified shortlist row for the mid-tier Claude option.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on SWE-bench Verified leaderboard (historical)