Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Shared

Benchmarks

Task NameDataset NameSOTA ResultTrend
LLM EvaluationShared (evaluation)
Tie-aware Accuracy78
10
Multi-judge evaluationShared 500-prompt sample
Global Correlation (r)0.87
5
Scientific Discovery Pairwise PreferenceShared 40-task (evaluation)
Win Count40
4
Humanoid Loco-Manipulation GenerationShared 20-object (test)
Affordance Realism74.7
4
Calibration and DiscriminationShared pooled aggregation (test)
Brier Score (BS)0.1
4
Showing 5 of 5 rows