cursor.com · benchmark source
Cursor
Cursor’s first-party agent evals for IDE-native multi-file work. CursorBench is vendor-run from real Cursor sessions — useful inside Cursor, not a public reproducible harness.
Live traffic
Online evals
Controlled experiments on live Cursor traffic that track interaction and output-quality proxies — used to catch regressions offline suites miss.
Why it matters
Developer experience can diverge from offline correctness; online evals keep product decisions grounded in real sessions.
Use it when
You are reading Cursor’s research posts about why a model default changed, and need the eval context.
The limitation
Online results are internal product metrics, not a public ranked table of models you can copy into a spreadsheet.
Board notes
Online evals — methodology notes
This board documents how the publisher evaluates — not a ranked shortlist. Open the source for the full write-up.
Local leader
Not transcribed
Use the live source
Rows shown
—
Online signal
Snapshot date
2026-08-07
8 days old · not a live API feed
Method
Controlled experiments on live Cursor traffic
Purpose
Catch regressions offline suites miss
Public table
Not published as a ranked model list
Read with
Offline CursorBench correctness + your own repo trials
The live source owns the complete table
Online evals are product metrics inside Cursor, not a downloadable leaderboard. Use this board for framing when Cursor explains a model-default change; use CursorBench for the public correctness shortlist.
Online evals are product metrics inside Cursor, not a downloadable leaderboard. Use this board for framing when Cursor explains a model-default change; use CursorBench for the public correctness shortlist.
Open CursorBench research postEvidence in this catalog
Where Online evals fits
Closest verified examples we currently carry. A missing score is not a zero.
Model examples
No directly comparable model score is verified here yet.
Harness examples
No directly comparable harness score is verified here yet.