Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Human Evaluation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Human Evaluation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
ESConv
KEMI
Win Rate
70
36
3mo ago
VOVILLE
PAVE
Score
6.18
9
2mo ago
NeurIPS ICML ICLR Proposals 2025 (test)
Stepwise CoT
Wins
31
6
3mo ago
CulturalVQA OOD (test)
MMBoundary
Faithfulness
7.66
6
4mo ago
ScienceVQA (test)
MMBoundary
Faithfulness Score
8.35
6
4mo ago
A-OKVQA (test)
MMBoundary
Faithfulness Score
7.83
6
4mo ago
Skywork (test)
AAD
Elo Rating
1,610.1
5
1mo ago
DeepResearch Bench 20 reports (sampled)
PTAH
Readability (Win/Tie Rate)
95
5
1mo ago
UltraFeedback 50 sampled questions
OTPO
Win Rate (Expert 1)
62
5
4mo ago
CoVOMIX2-DIALOGUE-20S and CoVOMIX2-DIALOGUE-WILDREF mix
SCENA
Win Rate (SCENA Preferred)
84.6
4
1mo ago
RCC-PVD (evaluation)
Ranking
Rank Preference Rate
61
4
4mo ago
Human Evaluation
Ann Brown
Trustworthiness
0.86
4
4mo ago
Human study blinded triplet comparison
Clean
Consistency Rank
1.5
3
22d ago
Human Evaluation Rapport and UX
IPA
Rapport Score R1
3.88
3
1mo ago
Management subset
MENTOR
Win Rate
93
3
1mo ago
Finance
MENTOR
Win Rate
97
3
1mo ago
Education subset
MENTOR
Win Rate
85
3
1mo ago
Human Evaluation Evil Players
GRAIL Agent
Contributed Success
3.78
3
3mo ago
MathQA
Ours
Accuracy
89.2
3
4mo ago
50 randomly selected model responses
GPT-4.1
Clarity
98
3
4mo ago
Human Evaluation Set (test)
LongDPO
Win Rate
0.65
3
4mo ago
200 human-generated instructions
Olympus
Success Rate
0.865
3
4mo ago
HH dataset
RRHF_DP
Win Rate
59
3
4mo ago
MS MARCO (test)
RBG
Preference: FiD
18
3
4mo ago
Qwen-Image rollout results
HPSv3++
Win Rate
77.5
2
1mo ago
Showing 25 of 33 rows
25 / page
50 / page
100 / page
1
2
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs