Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Agentic on Terminal Bench 2.0
Loading...
48.31
Pass@1
GLM-5
-0.7884
11.9583
24.705
37.4517
Mar 19, 2026
Apr 6, 2026
Apr 24, 2026
May 12, 2026
May 30, 2026
Jun 17, 2026
Jul 6, 2026
Pass@1
Updated 18d ago
Evaluation Results
Method
Method
Links
Pass@1
GLM-5
Thinking Mode=non-thin...
2026.06
48.31
Kimi-K2.5
Mode=Instant
2026.06
48.3
GPT-5.4
Reasoning Mode=non-rea...
2026.06
46.07
Qwen3.5 35B-A3B
2026.03
40.5
Ling-2.6-1T
2026.06
40.45
Nemotron-3-Super 120B-A12B
2026.03
31
DeepSeek-V3.2
Thinking Mode=nothink
2026.06
29.21
Nemotron-Cascade-2 30B-A3B
2026.03
21.1
Audex 30B-A3B
Context Length=1M
2026.07
19.1
Nemotron-3-Nano 30B-A3B
2026.03
8.5
Qwen3-Omni 30B-A3B Thinking
Context Length=64K
2026.07
1.1
Feedback
Search any
task
Search any
task