Skip to main content
← Benchmark sources

arena.ai · benchmark source

Arena

Live human-preference and agent boards from Arena (LMArena). Scores are board-specific — Text Arena Score, WebDev Arena Score, and Agent net improvement are not interchangeable.

Open arena.ai3 boards on this publisher

Web development

Code Arena · WebDev

A live human-preference leaderboard for models that turn natural-language prompts into working front-end web applications, scored with Arena’s Bradley-Terry Arena Score.

Why it matters

It captures the prompt-to-product loop — layout, interactions, and whether the app feels usable — rather than only repository issue repair.

Use it when

You are choosing a model for UI/app generation and want a live preference signal alongside coding-agent benchmarks.

The limitation

Votes reflect voter taste on sandboxed app demos. Arena Score is board-specific and not comparable to Agent net improvement, SWE-bench, or DeepSWE percentages.

Leaderboard snapshot

Code Arena · WebDev results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

Full Arena WebDev board mirrored — not a catalog-only shortlist.

Local leader

claude-opus-5-max

1,691

Rows shown

115

Arena Score

Snapshot date

2026-08-12

3 days old · not a live API feed

Score profile

Arena Score by published row

Higher is better in this view

claude-opus-5-max
1,691
kimi-k3-max
1,674
qwen3.8-max
1,669
claude-opus-5-high
1,664
grok-4.6-high
1,630
claude-fable-5
1,627
gpt-5.6-sol-xhigh (codex-harness)
1,622
gemini-3.7-flash-high
1,588
glm-5.2-max
1,587
deepseek-v4-flash-high
1,582
claude-opus-4-8-thinking
1,564
claude-opus-4-7
1,558
claude-opus-4-7-thinking
1,557
grok-4.5
1,553
claude-opus-4-6-thinking
1,545
claude-sonnet-5-high
1,541
claude-opus-4-8
1,539
muse-spark-1.1
1,539
claude-opus-4-6
1,537
gemini-3.6-flash
1,537
muse-spark-1.2 (xHigh)
1,535
claude-sonnet-4-6
1,524
gpt-5.6-terra-xhigh (codex-harness)
1,524
hy3
1,524
seed-2.1-pro-preview
1,523
gpt-5.6-luna-xhigh (codex-harness)
1,519
qwen3.7-max-20260517
1,517
glm-5.1
1,512
gpt-5.5-xhigh (codex-harness)
1,509
kimi-k2.6
1,509
gemini-3.5-flash-high
1,506
claude-opus-4-5-20251101-thinking-32k
1,495
gemini-3.5-flash
1,492
minimax-m3
1,491
gemini-3.5-flash-medium
1,487
gpt-5.5-high (codex-harness)
1,486
qwen3.6-max-preview
1,479
kimi-k2.7-code
1,474
mimo-v2.5-pro
1,474
claude-opus-4-5-20251101
1,468
deepseek-v4-pro-high-preview
1,464
gpt-5.4-high (codex-harness)
1,463
qwen3.6-plus
1,459
gpt-5.5 (codex-harness)
1,458
gemini-3.5-flash-lite
1,450
gemini-3.1-pro-preview
1,447
deepseek-v4-pro
1,445
gpt-5.4-medium (codex-harness)
1,442
gemini-3-flash
1,438
gemini-3-pro
1,438
mimo-v2.5
1,438
glm-5
1,436
kimi-k2.5-thinking
1,436
glm-4.7
1,434
mimo-v2-pro
1,434
deepseek-v4-flash-high-preview
1,431
gpt-5-medium
1,419
gpt-5.2
1,418
gpt-5.3-codex (codex-harness)
1,409
inkling
1,406
kimi-k2.5-instant
1,405
glm-5v-turbo
1,400
qwen3.5-397b-a17b
1,400
gpt-5.4-mini-high
1,398
minimax-m2.7
1,398
claude-sonnet-4-5-20250929-thinking-32k
1,392
gpt-5.1-medium
1,391
gpt-5.4
1,390
claude-opus-4-1-20250805
1,389
minimax-m2.1-preview
1,387
claude-sonnet-4-5-20250929
1,385
minimax-m2.5
1,384
gemini-3-flash (thinking-minimal)
1,383
grok-4.20-beta-0309-reasoning
1,374
solar-pro4
1,373
gpt-5.3-codex (codex-harness)
1,370
gemma-4-31b
1,366
gemma-4-26b-a4b
1,362
deepseek-v3.2-thinking
1,360
muse-glimmer
1,359
qwen3.5-122b-a10b
1,358
qwen3.5-27b
1,357
hunyuan-hy3-preview
1,356
grok-4.3
1,355
laguna-m.1
1,348
gpt-5.1
1,341
glm-4.6
1,340
gpt-5.2-codex
1,338
gpt-5.1-codex
1,336
mimo-v2-flash (non-thinking)
1,330
claude-haiku-4-5-20251001
1,326
deepseek-v3.2
1,324
kimi-k2-thinking-turbo
1,322
laguna-xs.2
1,303
minimax-m2
1,297
mimo-v2-flash (thinking)
1,292
qwen3-coder-480b-a35b-instruct
1,273
deepseek-v3.2-exp
1,272
mistral-medium-3.5
1,266
KAT-Coder-Pro-V1
1,255
gemini-3.1-flash-lite-preview
1,254
qwen3.5-35b-a3b
1,250
gpt-5.1-codex-mini
1,243
grok-4-1-fast-reasoning
1,240
trinity-large-thinking
1,239
qwen3.5-flash
1,238
mistral-large-3
1,230
gemini-2.5-pro
1,225
grok-4.1-thinking
1,210
devstral-2
1,194
granite-4.1-8b
1,193
mercury-2
1,166
grok-code-fast-1
1,164
grok-4-fast-reasoning
1,162
devstral-medium-2507
1,080

Full Code Arena WebDev overall board mirrored 2026-08-12 from arena.ai/leaderboard/code/webdev. Arena Score is not a percentage and is not comparable to Agent net improvement or DeepSWE pass rates. Open the live source for confidence intervals and preliminary tags.

Open Code Arena · WebDev leaderboard

Evidence in this catalog

Where Code Arena · WebDev fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

  • Claude Opus 51691 Arena Score (max)

    WebDev overall board retrieved 2026-08-12.

  • Kimi K31674 Arena Score (max)

    WebDev overall board retrieved 2026-08-12.

  • GPT-5.6 Sol1622 Arena Score (xhigh · codex)

    Codex-harness row on the WebDev board retrieved 2026-08-12.

Harness examples

No directly comparable harness score is verified here yet.

Open this board on Code Arena · WebDev leaderboard