Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM Judging Agreement on Reasoning Agreement Benchmark 500-sample human-annotated
Loading...
0.91
Cohen's Kappa
Multi-sub RM
0.0676
0.2863
0.505
0.7237
Jul 11, 2025
Cohen's Kappa
Updated 1mo ago
Evaluation Results
Method
Method
Links
Cohen's Kappa
Multi-sub RM
2025.07
0.91
GPT-4o
2025.07
0.9
Master-RM-7B
Model Scale=7B
2025.07
0.9
Qwen2.5-72B-Instruct
Model Scale=72B
2025.07
0.88
Qwen2.5-32B-Instruct
Model Scale=32B
2025.07
0.88
Qwen2.5-14B-Instruct
Model Scale=14B
2025.07
0.88
Master-RM-32B
Model Scale=32B
2025.07
0.87
Qwen2.5-1.5B-Instruct
Model Scale=1.5B
2025.07
0.83
Qwen2.5-3B-Instruct
Model Scale=3B
2025.07
0.82
Omni-Judge
2025.07
0.81
LLaMA3-70B-Instruct
Model Scale=70B
2025.07
0.81
Qwen2.5-7B-Instruct
Model Scale=7B
2025.07
0.8
LLaMA3-8B-Instruct
Model Scale=8B
2025.07
0.73
General-Verifier
2025.07
0.7
Qwen2.5-0.5B-Instruct
Model Scale=0.5B
2025.07
0.1
Feedback
Search any
task
Search any
task