Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Puzzle Solving on PUZZLE Human hard
Loading...
5.7
Accuracy
GPT-5
-0.228
1.311
2.85
4.389
Jun 2, 2026
Accuracy
Mean Completion Tokens
Updated 1mo ago
Evaluation Results
Method
Method
Links
Accuracy
Mean Completion Tokens
GPT-5
2026.06
5.7
19,861.9
Kimi K2
2026.06
1
61,307.1
Qwen 3 235B
Parameters=235B
2026.06
0
23,608.6
DeepSeek V3.2
2026.06
0
36,787.4
Feedback
Search any
task
Search any
task