Skip to main content
← Benchmark sources

artificialanalysis.ai · benchmark source

Artificial Analysis

Independent evaluations across composite intelligence, science reasoning, and agentic terminal tasks. Multiple boards live under one publisher.

Open artificialanalysis.ai3 boards on this publisher

Tool use

Terminal-Bench v2.1

A terminal-oriented evaluation of agents that must use shell tools and complete multi-step tasks in an environment.

Why it matters

It tests the loop around a model: planning, commands, file changes, and validation.

Use it when

You are choosing a terminal agent and want an independent harness signal about multi-step tool interaction.

The limitation

A high score does not guarantee safe behaviour in your repository. Local rows here use Artificial Analysis Terminus 2; proprietary agent harnesses on tbench.ai are separate.

Leaderboard snapshot

Terminal-Bench v2.1 results in a readable view

Use the tabs for the primary score and other published fields from this board. The source link remains authoritative for the live table.

Full Artificial Analysis Terminus 2 table mirrored here. Proprietary agent harnesses on tbench.ai are not mixed into this board.

Local leader

GPT-5.6 Sol (xhigh)

89.5%

Rows shown

191

Success rate

Snapshot date

2026-08-14

1 day old · not a live API feed

Score profile

Success rate · Artificial Analysis Terminus 2

Higher is better in this view

GPT-5.6 Sol (xhigh)
89.5%
Claude Opus 5 (Adaptive Reasoning, Max Effort)
89.1%
Grok 4.6 (high)
88.4%
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
88%
GPT-5.6 Sol (max)
88%
GPT-5.6 Terra (max)
88%
Claude Opus 5 (Adaptive Reasoning, High Effort)
87.6%
GPT-5.6 Sol (high)
87.3%
Claude Opus 5 (Adaptive Reasoning, Medium Effort)
86.1%
GPT-5.6 Sol (medium)
86.1%
Gemini 3.7 Flash (high)
85.8%
Kimi K3 (max)
85%
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
84.6%
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
84.6%
GPT-5.5 (xhigh)
84.3%
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
83.1%
Kimi K3 (low)
82.4%
Grok 4.5 (high)
81.6%
Qwen3.8 Max
81.3%
GPT-5.6 Luna (max)
80.9%
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
80.5%
GPT-5.5 (medium)
80.5%
GPT-5.6 Terra (xhigh)
80.1%
Muse Spark 1.2 (xhigh)
80.1%
GPT-5.5 (high)
79.4%
DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
78.7%
DeepSeek V4 Pro 0813 (max)
78.7%
Gemini 3.5 Flash (high)
78.7%
GPT-5.4 (xhigh)
78.3%
GLM-5.2 (max)
77.9%
GPT-5.6 Luna (xhigh)
77.9%
Muse Spark 1.1 (xhigh)
77.9%
Gemini 3.6 Flash (high)
77.5%
GPT-5.6 Sol (low)
76.8%
Claude Opus 5 (Adaptive Reasoning, Low Effort)
76.4%
GPT-5.6 Terra (high)
75.7%
Claude Sonnet 5 (Non-reasoning, High Effort)
75.3%
Qwen3.7 Max
74.5%
GPT-5.6 Sol (Non-reasoning)
74.2%
Gemini 3.1 Pro Preview
73.8%
GPT-5.6 Terra (medium)
72.3%
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
71.2%
Motif 3 (Beta)
70.8%
KAT Coder Pro V2
70%
GPT-5.6 Luna (high)
69.7%
Nex-N2-Pro
67.8%
Kimi K2.7 Code
67.4%
Agnes 2.5 Pro Alpha
67%
Kimi K2.6
65.9%
GPT-5.5 (low)
65.5%
MiMo-V2.5-Pro
65.2%
MiniMax-M3
65.2%
DeepSeek V4 Pro (Reasoning, High Effort)
64.8%
Hy3
64.4%
MiMo-V2.5
63.7%
GPT-5.6 Terra (low)
62.5%
Muse Spark
62.2%
DeepSeek V4 Flash (Reasoning, Max Effort)
61.8%
GLM-5.1 (Reasoning)
61.8%
MiMo-V2-Flash (Non-reasoning)
61.8%
Qwen3.6 Plus
61.4%
GPT-5.5 (Non-reasoning)
61%
Qwen3.7 Plus
61%
GPT-5.4 nano (xhigh)
60.7%
Qwen3.6 27B (Reasoning)
60.7%
JT-4.1 Flash 236B A21B
59.6%
GPT-5.4 mini (xhigh)
59.2%
DeepSeek V4 Flash (Reasoning, High Effort)
56.9%
GPT-5.6 Terra (Non-reasoning)
56.2%
Claude 4.5 Sonnet (Reasoning)
55.8%
Ling 3.0 Flash
55.4%
MiniMax-M2.7
55.4%
Inkling (xhigh)
55.1%
Inkling Small
55.1%
Nemotron 3 Ultra 550B A55B (Reasoning)
53.9%
Gemini 3.5 Flash-Lite
53.6%
GPT-5.6 Luna (medium)
53.2%
GPT-5.1 (high)
52.4%
Grok Build 0.1 0616
52.1%
GLM-5.2 (Non-reasoning)
51.7%
Muse Glimmer (high)
51.7%
Qwen3.5 397B A17B (Reasoning)
51.3%
Qwen3.6 27B (Non-reasoning)
51.3%
Mistral Medium 3.5
50.6%
LongCat 2.0
50.2%
GLM-4.6 (Reasoning)
49.4%
Qwen3.5 122B A10B (Reasoning)
47.6%
Qwen3.5 122B A10B (Non-reasoning)
47.2%
DeepSeek V3.2 (Reasoning)
46.8%
Kimi K2.5 (Reasoning)
45.7%
GLM-4.7 (Reasoning)
45.3%
DeepSeek V3.1 Terminus (Reasoning)
44.9%
Qwen3.6 35B A3B (Reasoning)
44.9%
Claude 4.5 Haiku (Reasoning)
44.2%
Gemma 4 31B (Reasoning)
43.4%
GPT-5.6 Luna (low)
43.4%
Ring-2.6-1T
43.1%
Qwen3.6 35B A3B (Non-reasoning)
41.6%
Qwen3.5 35B A3B (Non-reasoning)
40.8%
Grok 4.3 (high)
39.7%
Step 3.7 Flash
39.3%
Gemma 4 26B A4B (Reasoning)
39%
GPT-5.6 Luna (Non-reasoning)
39%
Nemotron 3 Super 120B A12B (Reasoning)
38.6%
Qwen3 Coder Next
38.2%
Claude 4 Sonnet (Reasoning)
36.3%
North Mini Code
35.6%
GPT-5 (high)
35.2%
GPT-5.5 Instant (June 2026)
34.8%
Grok 4.3 (Non-reasoning)
34.1%
Gemini 3.1 Flash-Lite
31.1%
Devstral 2
30.3%
K-EXAONE (Reasoning)
30.3%
Devstral Small 2
29.6%
Nova 2.0 Pro Preview (medium)
29.6%
Gemma 4 31B (Non-reasoning)
29.2%
Qwen3.5 9B (Reasoning)
29.2%
G9v3-39A5B
28.5%
Gemini 2.5 Pro
28.5%
Ling 3.0 Tiny
27.7%
Gemma 4 12B (Reasoning)
27.3%
Mercury 2
27.3%
gpt-oss-120b (high)
26.2%
Mistral Small 3.1
26.2%
Qwen3.5 4B (Reasoning)
25.8%
Ling 2.6 Flash
24.3%
Nemotron 3.5 Lightning
24.3%
Command A+
22.8%
EXAONE 4.5 33B
21.3%
Qwen3.5 4B (Non-reasoning)
21.3%
Qwen3.5 9B (Non-reasoning)
21.3%
Mistral Small 4 (Reasoning)
21%
Nemotron Cascade 2 30B A3B
20.6%
Trinity Large Thinking
20.6%
Nova 2.0 Pro Preview (low)
19.5%
DeepSeek R1 (Jan '25)
19.1%
HyperNova 60B 2605
18.4%
Nova 2.0 Pro Preview (Non-reasoning)
17.2%
DeepSeek V3 (Dec '24)
16.9%
Nova 2.0 Lite (high)
16.1%
K2 Think V2
15%
DeepSeek V3 0324
13.9%
gpt-oss-120b (low)
13.9%
gpt-oss-20b (high)
13.9%
Mistral Medium 3.1
13.9%
DiffusionGemma 26B A4B
12.4%
Magistral Medium 1.2
12.4%
Mistral Large 3
12%
Qwen3 235B A22B 2507 (Reasoning)
12%
Solar Pro 3
12%
Celeris-1
11.2%
Claude 3.5 Haiku
10.1%
GPT-4.1 mini
10.1%
Ministral 3 14B
9.7%
Llama 4 Maverick
7.9%
Nemotron 3 Nano Omni 30B A3B Reasoning
6.7%
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)
6.7%
Qwen3 Next 80B A3B (Reasoning)
6.7%
G9v3-3B
6%
GPT-4o mini
5.6%
Mistral Small 3.2
5.6%
Qwen3 32B (Reasoning)
5.2%
Llama 3.3 Instruct 70B
4.9%
Qwen3 14B (Reasoning)
4.9%
Gemma 3 27B Instruct
4.5%
Magistral Small 1.2
4.5%
o3-mini (high)
4.5%
Ministral 3 8B
4.1%
GPT-4.1 nano
3.7%
GPT-5 mini (high)
3.7%
Llama 4 Scout
3.7%
NVIDIA Nemotron 3 Nano 4B
3.7%
Granite 4.1 8B
3.4%
Qwen3.5 2B (Reasoning)
3%
Granite 4.1 30B
2.6%
Qwen3 8B (Reasoning)
2.2%
Gemma 4 E4B (Reasoning)
1.9%
Llama 3.1 Instruct 8B
1.5%
Qwen3 30B A3B 2507 (Reasoning)
1.5%
Granite 4.1 3B
1.1%
Nanbeige4.1-3B
1.1%
Gemma 3n E4B Instruct
0.7%
Gemma 3 4B Instruct
0.4%
Gemma 4 E2B (Reasoning)
0.4%
Phi-4 Mini Instruct
0.4%
Qwen3.5 0.8B (Non-reasoning)
0.4%
Gemma 3 12B Instruct
0%
MiniCPM-V 4.6 1.3B
0%
Ministral 3 3B
0%
Qwen3.5 0.8B (Reasoning)
0%
Qwen3.5 2B (Non-reasoning)
0%

Full Artificial Analysis Terminus 2 / Terminal-Bench v2.1 table mirrored 2026-08-11 from artificialanalysis.ai (Pass@1, 3 repeats), with Grok 4.6 (high) 88.4% inserted from the live AA table on 2026-08-14. Proprietary agent harnesses published on tbench.ai are intentionally excluded so scores stay on one evaluator.

Open Artificial Analysis Terminal-Bench v2.1

Evidence in this catalog

Where Terminal-Bench v2.1 fits

Closest verified examples we currently carry. A missing score is not a zero.

Model examples

  • GPT-5.6 Sol89.5% (xhigh · AA Terminus 2)

    Artificial Analysis Terminus 2 pass@1 on Terminal-Bench v2.1.

  • Claude Opus 589.1% (max · AA Terminus 2)

    Artificial Analysis Terminus 2 pass@1 on Terminal-Bench v2.1.

  • Claude Fable 584.6% (AA Terminus 2)

    Artificial Analysis Terminus 2 pass@1 on Terminal-Bench v2.1 — not a separate provider-only figure.

Harness examples

  • Claude CodeTerminal agent

    Terminal-Bench scores models inside a harness loop — compare terminal agents after you narrow the model field.

  • CursorIDE agent harness

    Use harness comparisons to separate editor workflow from bare API scores on this board.

Open this board on Artificial Analysis Terminal-Bench v2.1