Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LLMBar

Benchmarks

Task NameDataset NameSOTA ResultTrend
Robustness EvaluationLLMBar
Accuracy83.07
8
LLM-as-a-Judge CalibrationLLMBar (test)
Test Risk (MSE)0.194
7
Reward ModelingLLMBar (test)
Test MSE (Table)0.2039
5
Quality-judgment accuracyLLMBar ZH condition
Strict Accuracy (%)86.5
4
Quality-judgment accuracyLLMBar EN condition
Strict Accuracy (LLMBar EN)89.5
4
Quality-judgment accuracyLLMBar LS
Strict Accuracy86.5
4
Showing 6 of 6 rows