Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

HH

Benchmarks

Task NameDataset NameSOTA ResultTrend
Value AlignmentHH Balance-8
Conformity Score4.317
17
Dialogue Preference LearningHH (test)
Win Rate (0% Flip)90.8
14
Human Preference AlignmentHH (test)
Reward3.8764
14
Response GenerationHH dataset
Reward-0.96
13
DialogueHH (Anthropic Helpful and Harmless)
Win Rate (0% Flip)82.5
10
Harmfulness EvaluationHH Harmless
Beaver-7B Cost Score3.25
10
AlignmentHH IDN 40%
Win Rate68
8
AlignmentHH (IDN 20%)
Win Rate78.2
8
Preference EvaluationHH-Helpful
Win Count52
8
Model DiscoveryHH
Avg NLL (Model)25.18
6
LLM-as-Judge evaluationHH dataset
WCWR59.1
5
Closed Loop RLHFHH (40% noise)
Win Rate57
3
Closed Loop RLHFHH 20% noise
Win Rate58.7
3
Closed Loop RLHFHH 0% noise
Win Rate61.6
3
Human EvaluationHH dataset
Win Rate59
3
LLM Preference Alignment EvaluationHH Helpful
Preference (spec vs ctl)55
1
Pairwise Judge ComparisonHH helpful
Win/Loss Count149
1
Showing 17 of 17 rows