Skip to main content
← Benchmark sources

arena.ai · benchmark source

Arena

Live human-preference and agent boards from Arena (LMArena). Scores are board-specific — Text Arena Score, WebDev Arena Score, and Agent net improvement are not interchangeable.

Open arena.ai3 boards on this publisher

Agent systems

Agent net improvement

A net-improvement percentage reported for agent systems on the Arena Agent leaderboard.

Why it matters

It provides an agent-specific external signal without presenting a percentage improvement as an Elo rating.

Use it when

You are comparing agent-system results on the Arena Agent board and understand the metric is not a general model rating.

The limitation

It reflects the Arena Agent board and evaluation mix, not factual accuracy, cost, or enterprise safety.

Leaderboard snapshot

Agent net improvement results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

Full Arena Agent board mirrored. Negative net improvement means the agent hurt outcomes versus the baseline on this board.

Local leader

Claude Opus 5 (High)

+12.19%

Rows shown

48

Net improvement

Snapshot date

2026-08-14

1 day old · not a live API feed

Score profile

Net improvement by published row

Higher is better in this view

Claude Opus 5 (High)
+12.19%
Claude Fable 5 (High)
+12.01%
Claude Opus 5 (Max)
+11.95%
GPT 5.6 Sol (xHigh)
+10.72%
Kimi K3 (Max)
+10.43%
Claude Opus 4.8 (Thinking)
+9.54%
GPT 5.5 (xHigh)
+8.70%
Claude Opus 4.7 (Thinking)
+8.17%
Claude Opus 4.7
+7.63%
GPT 5.5 (High)
+7.63%
Claude Sonnet 5 (High)
+7.37%
Claude Opus 4.6
+6.70%
GLM 5.2 (Max)
+6.67%
GPT 5.5
+6.23%
Grok 4.5
+5.73%
GPT 5.4 (High)
+4.99%
GPT 5.6 Luna (xHigh)
+4.29%
Deepseek V4 Flash (High) (20260731)
+3.93%
Gemini 3.7 Flash (High)
+3.61%
GPT 5.6 Terra (xHigh)
+3.43%
Claude Sonnet 4.6
+2.99%
Claude Opus 4.8
+2.50%
Muse Spark 1.1
+1.09%
Kimi K2.7 Code
+1.03%
GLM 5.1
+0.53%
Qwen3.7 Max
-0.01%
DeepSeek V4 Pro
-0.07%
Gemini 3.5 Flash (High)
-0.43%
Gemini 3.1 Pro Preview
-0.58%
Kimi K2.6
-0.63%
Hy3
-1.30%
Qwen3.7 Plus
-1.81%
Mimo V2.5 Pro
-2.21%
DeepSeek V4 Flash
-2.33%
Gemini 3.6 Flash
-2.36%
Minimax M3
-2.52%
Gemini 3.5 Flash (Medium)
-3.64%
Inkling
-6.71%
Mistral Medium 3.5
-6.93%
Grok 4.3 (High)
-8.47%
Gemini 3 Flash
-8.59%
Grok Build 0.1
-9.04%
Gemini 3.5 Flash-Lite
-10.21%
Minimax M2.7
-11.13%
Solar Pro 4
-12.10%
Nemotron 3 Ultra
-14.56%
Grok 4.3
-14.61%
Gemma 4 31B
-18.23%

Full Arena Agent net-improvement board mirrored from arena.ai/leaderboard/agent. Gemini 3.7 Flash (High) is the live net-improvement column (3.61% as of 2026-08-14), not the success-rate column (~9.9%). The live board updates continuously; confidence intervals often overlap below the top cluster, so treat mid-board order as directional.

Open Arena Agent leaderboard

Evidence in this catalog

Where Agent net improvement fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

  • Claude Opus 5+12.19% net improvement (high)

    Arena Agent snapshot retrieved 2026-08-14; mid-board order can move as sessions accumulate.

  • Claude Fable 5+12.01% net improvement (high)

    Arena Agent snapshot retrieved 2026-08-14.

  • GPT-5.6 Sol+10.72% net improvement (xhigh)

    Arena Agent snapshot retrieved 2026-08-14.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on Arena Agent leaderboard