Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hard Reasoning on HLE
Loading...
37.7
Pass@1
Gemini-3.0 Pro
5.772
14.061
22.35
30.639
Dec 2, 2025
Jan 2, 2026
Feb 3, 2026
Mar 6, 2026
Apr 7, 2026
May 8, 2026
Jun 9, 2026
Pass@1
Updated 1mo ago
Evaluation Results
Method
Method
Links
Pass@1
Gemini-3.0 Pro
2025.12
37.7
GPT-5 High
2025.12
26.3
DeepSeek-V3.2
thinking mode=true, te...
2025.12
25.1
Kimi-K2
thinking mode=true
2025.12
23.9
DeepSeek-V3.2
thinking mode=true, te...
2025.12
23.9
Claude-4.5-Sonnet
2025.12
13.7
MiniMax M2
2025.12
12.5
Qwen3 + SFT + B-DAPO w/ recycling
Size=4B, k=2, Recycle=...
2026.06
10.4
SmartSearch-3B
Size=3B
2026.06
9.8
Qwen3 + SFT + B-DAPO w/ recycling
Size=1.7B, k=2, Recycl...
2026.06
8
BehaviorPrime-1.7B
Size=1.7B
2026.06
7.8
ASearcher-Web-7B
Size=7B
2026.06
7
Feedback
Search any
task
Search any
task