Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA LLM-as-a-Judge benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
LLM-as-a-Judge
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
PreferenceBench
CalibraEval
Accuracy
90.71
59
3mo ago
MTbench (test)
DI
StdDev
2.24
45
4mo ago
MT-Bench
PA-GRPO
Accuracy
81.4
44
3mo ago
RewardBench 1.0 (test)
CC
Rstd
0.54
36
4mo ago
RewardBench
Qwen3-Next-80B-A3B-Thinking
Accuracy
92.9
31
3mo ago
JudgeBench
DeepSeek-V3
Accuracy
84.19
29
4mo ago
Peer-Support Evaluation Set
MindTailor
Empathy
4.89
23
1mo ago
PreferenceBench
PA-GRPO
Accuracy
90.2
21
4mo ago
High-contrast response pairs
LongCat-Flash-Chat
Discriminability (πi)
0.87
20
2mo ago
ARENA
EpiPersona-A
Accuracy
66.07
20
3mo ago
PRISM
EpiPersona-A
Accuracy
59.38
20
3mo ago
SenseBench
SenseJudge
Math Score
86.53
17
1mo ago
PRISM (test)
SynthesizeMe
Accuracy
58.9
14
4mo ago
Chatbot Arena (test)
Gemini-2.5-Pro
Accuracy
68.13
14
4mo ago
FairJudge Benchmark 1K (test)
FairJudge-8B
Agreement
71.5
13
4mo ago
JudgeLM (test)
Qwen2.5-72B
Agreement
79.59
13
4mo ago
PandaLM Human Annotations (test)
FairJudge-8B
Agreement
0.7683
13
4mo ago
TL;DR
Heuristic Selection
Coverage
82.6
12
2mo ago
Chatbot Arena
Heuristic Selection
Coverage
94.3
12
2mo ago
HH-RLHF
Heuristic Selection
Coverage
81.3
12
2mo ago
AlpacaEval
Heuristic Selection
Coverage
78.3
12
2mo ago
Preference Bench (test)
CalibraEval
Std Dev
2.82
9
4mo ago
RewardBench (test)
CalibraEval
Std Dev (Reward)
2.72
9
4mo ago
LLM-as-a-Judge (10-fold cross-validation)
Qwen3-14B
CG Accuracy
88
8
1mo ago
JudgeBench (Merged GPT Claude)
qwen3.5-35b
Direct Baseline Score
87.38
8
3mo ago
Showing 25 of 30 rows
25 / page
50 / page
100 / page
1
2
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs