cursor.com · benchmark source
Cursor
Cursor’s first-party agent evals for IDE-native multi-file work. CursorBench is vendor-run from real Cursor sessions — useful inside Cursor, not a public reproducible harness.
Coding agents
CursorBench
Cursor’s offline correctness suite for ambiguous, multi-file coding-agent tasks sourced from real Cursor sessions rather than public GitHub issues.
Why it matters
It separates frontier models on IDE-style agent work where SWE-bench and similar public suites are increasingly saturated or contaminated.
Use it when
You are choosing a model for Cursor’s agent loop and want a signal aligned with Cursor’s own product evals — then validate on your repo.
The limitation
First-party and not independently reproducible from a public harness. Scores move when Cursor refreshes the problem distribution (compare only within the same version).
Leaderboard snapshot
CursorBench — chart and full table
Same layout as Cursor’s public board: Cost / Tokens / Steps chart, then one full Score · Cost · Tokens · Steps table.
Local leader
Grok 4.6 Extra High
70.8%
Rows shown
56
Full CursorBench 3.2 table
Snapshot date
2026-08-14
Mirrored from cursor.com/cursorbench
Better is toward the top-right (higher score, lower cost). X-axis is reversed to match Cursor’s public chart.
| # | Model | Score | Cost / task | Tokens / task | Steps / task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 Extra High | 70.8% | $2.81 | 41,136 | 46 |
| 2 | Fable 5 Max | 70.5% | $17.32 | 103,525 | 72 |
| 3 | Opus 5 Max | 70.0% | $8.23 | 61,838 | 78 |
| 4 | Grok 4.6 High | 69.9% | $2.34 | 32,449 | 39 |
| 5 | Opus 5 Extra High | 69.3% | $7.35 | 54,239 | 72 |
| 6 | Fable 5 Extra High | 68.4% | $11.73 | 64,971 | 56 |
| 7 | GPT-5.6 Sol Max | 67.2% | $5.69 | 28,320 | 48 |
| 8 | Grok 4.6 Medium | 67.1% | $1.28 | 17,942 | 29 |
| 9 | Opus 5 High | 66.7% | $3.91 | 27,932 | 48 |
| 10 | Fable 5 High | 66.5% | $8.77 | 43,747 | 48 |
| 11 | Fable 5 Medium | 65.2% | $6.80 | 30,366 | 41 |
| 12 | GPT-5.6 Terra Max | 64.9% | $2.31 | 32,969 | 47 |
| 13 | GPT-5.6 Sol Extra High | 64.5% | $3.88 | 19,699 | 38 |
| 14 | Opus 5 Medium | 64.3% | $3.29 | 23,612 | 44 |
| 15 | GPT-5.6 Sol High | 63.5% | $2.79 | 13,867 | 32 |
| 16 | Opus 5 Low | 62.8% | $2.55 | 18,529 | 37 |
| 17 | Opus 4.8 Max | 62.3% | $5.77 | 71,411 | 44 |
| 18 | Fable 5 Low | 62.1% | $4.46 | 18,182 | 31 |
| 19 | Gemini 3.7 Flash High | 61.6% | $1.20 | 38,448 | 99 |
| 20 | Sonnet 5 Max | 61.5% | $4.30 | 92,882 | 86 |
| 21 | GPT-5.6 Luna Max | 61.1% | $0.39 | 87,973 | 61 |
| 22 | Grok 4.6 Low | 61.0% | $0.70 | 10,658 | 23 |
| 23 | Kimi K3 Max | 60.8% | $2.70 | 38,428 | 57 |
| 24 | GPT-5.6 Sol Medium | 60.0% | $1.95 | 9,747 | 27 |
| 25 | Kimi K3 High | 59.7% | $1.89 | 26,846 | 47 |
| 26 | Opus 4.8 Extra High | 59.4% | $4.50 | 51,121 | 40 |
| 27 | GPT-5.6 Terra Extra High | 59.2% | $1.15 | 16,089 | 29 |
| 28 | Gemini 3.7 Flash Medium | 59.0% | $0.95 | 30,953 | 82 |
| 29 | Sonnet 5 Extra High | 58.7% | $2.77 | 52,871 | 67 |
| 30 | GPT-5.5 Extra High | 58.4% | $2.85 | 17,534 | 32 |
| 31 | GPT-5.5 High | 58.4% | $2.05 | 12,183 | 28 |
| 32 | Opus 4.8 High | 58.0% | $3.15 | 33,548 | 33 |
| 33 | GPT-5.6 Luna Extra High | 57.7% | $0.23 | 22,480 | 48 |
| 34 | Sonnet 5 High | 56.9% | $2.13 | 39,483 | 57 |
| 35 | GPT-5.6 Luna High | 56.8% | $0.16 | 15,141 | 40 |
| 36 | Composer 2.5 | 56.1% | $0.44 | 14,286 | 33 |
| 37 | Opus 4.8 Medium | 56.1% | $2.81 | 28,384 | 32 |
| 38 | GLM 5.2 Max | 55.0% | $1.76 | 35,946 | 58 |
| 39 | GPT-5.6 Terra High | 54.2% | $0.71 | 9,468 | 23 |
| 40 | GPT-5.5 Medium | 53.8% | $1.51 | 8,522 | 25 |
| 41 | Gemini 3.7 Flash Low | 53.8% | $0.74 | 20,594 | 68 |
| 42 | Gemini 3.6 Flash High | 53.5% | $1.56 | 30,436 | 64 |
| 43 | Opus 4.8 Low | 53.1% | $2.02 | 19,624 | 27 |
| 44 | GPT-5.6 Sol Low | 52.6% | $1.01 | 5,104 | 19 |
| 45 | Sonnet 5 Medium | 52.4% | $1.44 | 26,200 | 46 |
| 46 | GLM 5.2 High | 51.5% | $1.19 | 21,829 | 49 |
| 47 | Gemini 3.6 Flash Medium | 51.2% | $1.48 | 28,511 | 62 |
| 48 | Kimi K3 Low | 50.5% | $0.99 | 13,007 | 33 |
| 49 | GPT-5.6 Terra Medium | 50.3% | $0.49 | 6,222 | 20 |
| 50 | Kimi K2.7 Code | 49.7% | $1.43 | 31,247 | 58 |
| 51 | GPT-5.6 Luna Medium | 47.7% | $0.08 | 7,095 | 28 |
| 52 | Sonnet 5 Low | 47.7% | $0.87 | 16,269 | 33 |
| 53 | Gemini 3.6 Flash Low | 47.4% | $1.13 | 20,529 | 50 |
| 54 | GPT-5.6 Terra Low | 46.9% | $0.42 | 5,312 | 19 |
| 55 | GPT-5.5 Low | 46.6% | $0.98 | 5,168 | 20 |
| 56 | GPT-5.6 Luna Low | 37.6% | $0.03 | 3,209 | 17 |
Avg cost / task applies each model’s published per-million-token pricing to tokens used on each task. This snapshot has no row-level training-data disclosure markers. Small score gaps may not be statistically meaningful — confirm on cursor.com/cursorbench.
Evidence in this catalog
Where CursorBench fits
Closest verified examples we currently carry. A missing score is not a zero.
Model examples
- Grok 4.6 Extra High70.8% · $2.81 / task
Top published CursorBench 3.2 row (full table mirrored 2026-08-12).
- Fable 5 Max70.5% · $17.32 / task
Within a point of Grok 4.6 Extra High — small gaps may not be meaningful.
- Opus 5 Max70.0% · $8.23 / task
Strong correctness at lower average cost than Fable 5 Max.
Harness examples
- CursorIDE agent harness
CursorBench scores models inside Cursor’s agent loop, not bare API chat.