cursor.com · benchmark source
Cursor
Cursor’s first-party agent evals for IDE-native multi-file work. CursorBench is vendor-run from real Cursor sessions — useful inside Cursor, not a public reproducible harness.
Methodology
Eval axes
Cursor evaluates agents on more than correctness: code quality, efficiency (including completion tokens), and interaction behaviour, plus controlled online traffic experiments.
Why it matters
A model can look strong on an offline grader and still feel worse in live use — Cursor’s online-offline loop is designed to catch that gap.
Use it when
You need the framing for how Cursor ships model defaults, not a second independent leaderboard.
The limitation
Most public snapshots only publish the correctness axis. Efficiency and interaction scores are described by Cursor but are not fully transcribed here as a ranked table.
Board notes
Eval axes — methodology notes
This board documents how the publisher evaluates — not a ranked shortlist. Open the source for the full write-up.
Local leader
Not transcribed
Use the live source
Rows shown
—
Eval axis
Snapshot date
2026-08-07
8 days old · not a live API feed
Correctness
Did the agent solve the task as graded offline?
Code quality
Is the change maintainable beyond a binary pass?
Efficiency
Tokens, steps, and cost relative to outcome
Interaction
How the agent behaves in the IDE loop with the developer
The live source owns the complete table
Cursor’s research post describes a multi-axis eval system. This page documents the axes; the ranked correctness table lives on the CursorBench board. Efficiency and interaction scores are not transcribed as a public leaderboard here.
Cursor’s research post describes a multi-axis eval system. This page documents the axes; the ranked correctness table lives on the CursorBench board. Efficiency and interaction scores are not transcribed as a public leaderboard here.
Open CursorBench research postEvidence in this catalog
Where Eval axes fits
Closest verified examples we currently carry. A missing score is not a zero.
Model examples
No directly comparable model score is verified here yet.
Harness examples
No directly comparable harness score is verified here yet.