Skip to main content

Benchmark sources

Start with the website that publishes the board

Each source below is a publisher. Open one to switch between the leaderboards they host — instead of treating every board as a separate site section.

cursor.com

IDE agent evals

Cursor

Cursor’s first-party agent evals for IDE-native multi-file work. CursorBench is vendor-run from real Cursor sessions — useful inside Cursor, not a public reproducible harness.

CursorBenchEval axesOnline evals

3 boards · 56 local rows

arena.ai

Arena

Live human-preference and agent boards from Arena (LMArena). Scores are board-specific — Text Arena Score, WebDev Arena Score, and Agent net improvement are not interchangeable.

Agent net improvementCode Arena · WebDevText Arena · Overall

3 boards · 553 local rows

artificialanalysis.ai

Artificial Analysis

Independent evaluations across composite intelligence, science reasoning, and agentic terminal tasks. Multiple boards live under one publisher.

Intelligence IndexGPQA DiamondTerminal-Bench v2.1

3 boards · 1342 local rows

deepswe.datacurve.ai

DeepSWE

Long-horizon engineering evaluation that publishes pass rate together with cost, output tokens, and agent steps.

DeepSWE 1.1

1 board · 24 local rows

swebench.com

SWE-bench

Repository-level GitHub issue resolution benchmarks. Verified is a human-filtered historical subset — treat frontier progress claims carefully.

SWE-bench Verified

1 board · 6 local rows

livebench.ai

LiveBench

A contamination-aware suite that refreshes questions over time. Category and subtask views live on the source site.

LiveBench overall

1 board · 43 local rows

Model providers

Provider catalogs

System limits from provider documentation — context capacity is not a quality benchmark.

Context window

1 board · 25 local rows

A better comparison habit

Use a board to create a shortlist, then test a task.

A publisher can host several boards. Stay on one source page and switch boards instead of mixing Arena Score with DeepSWE pass rates — or CursorBench correctness — as if they were the same unit.

Trust signal

Named publisher. Named board. Measurement date.

Open the source website when you need the full live table — our pages keep a readable local shortlist and the caveats that travel with it.

Example live write-up: CursorBench