Skip to main content
← Benchmark sources

cursor.com · benchmark source

Cursor

Cursor’s first-party agent evals for IDE-native multi-file work. CursorBench is vendor-run from real Cursor sessions — useful inside Cursor, not a public reproducible harness.

Open cursor.com3 boards on this publisher

Methodology

Eval axes

Cursor evaluates agents on more than correctness: code quality, efficiency (including completion tokens), and interaction behaviour, plus controlled online traffic experiments.

Why it matters

A model can look strong on an offline grader and still feel worse in live use — Cursor’s online-offline loop is designed to catch that gap.

Use it when

You need the framing for how Cursor ships model defaults, not a second independent leaderboard.

The limitation

Most public snapshots only publish the correctness axis. Efficiency and interaction scores are described by Cursor but are not fully transcribed here as a ranked table.

Board notes

Eval axes — methodology notes

This board documents how the publisher evaluates — not a ranked shortlist. Open the source for the full write-up.

Local leader

Not transcribed

Use the live source

Rows shown

Eval axis

Snapshot date

2026-08-07

8 days old · not a live API feed

Correctness

Did the agent solve the task as graded offline?

Code quality

Is the change maintainable beyond a binary pass?

Efficiency

Tokens, steps, and cost relative to outcome

Interaction

How the agent behaves in the IDE loop with the developer

The live source owns the complete table

Cursor’s research post describes a multi-axis eval system. This page documents the axes; the ranked correctness table lives on the CursorBench board. Efficiency and interaction scores are not transcribed as a public leaderboard here.

Cursor’s research post describes a multi-axis eval system. This page documents the axes; the ranked correctness table lives on the CursorBench board. Efficiency and interaction scores are not transcribed as a public leaderboard here.

Open CursorBench research post

Evidence in this catalog

Where Eval axes fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

No directly comparable model score is verified here yet.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on CursorBench research post