Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Puzzle Solving on PUZZLE Human normal
Loading...
58.1
Accuracy
GPT-5
19.516
29.533
39.55
49.567
Jun 2, 2026
Accuracy
Mean Completion Token Count
Updated 1mo ago
Evaluation Results
Method
Method
Links
Accuracy
Mean Completion Token Count
GPT-5
2026.06
58.1
17,273.6
DeepSeek V3.2
2026.06
44.8
27,037.7
Kimi K2
2026.06
41
43,751.2
Qwen 3 235B
Parameters=235B
2026.06
21
23,104
Feedback
Search any
task
Search any
task