Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LiveBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
ReasoningLiveBench Reasoning
Accuracy92
80
General ReasoningLiveBench
Accuracy53.47
55
ReasoningLiveBench
Accuracy31.1
40
Ensemble Committee SelectionLiveBench (test)
Mean θtest99.07
34
Code GenerationLiveBench (test)
Sig. Score56.5
26
General LLM BenchmarkingLiveBench
Official Score49.6
24
Code GenerationLiveBench
Avg@842.9
22
Code GenerationLiveBench
Signal58.7
21
General Language ModelingLiveBench
Accuracy31.1
17
Mathematical ReasoningLiveBench Math
Initial Task Score58.1
16
ReasoningLiveBench
Accuracy33
16
General EvaluationLiveBench
Accuracy46.83
15
CodingLiveBench
Accuracy40.23
15
Language Model EvaluationLiveBench
Average Score72.8
12
Code generationLiveBench
Accuracy31.1
12
Mathematical ReasoningLiveBench
Accuracy53.6
12
Code generationLiveBench
Pass@1020.39
8
Single-event Scene Revisit (Different Pose)LiveBench
DINO Feature Similarity (FG)0.691
8
Single-event Scene Revisit (Same Pose)LiveBench
PSNR (Background)20.132
8
Instruction FollowingLiveBench IFEval
Quality100
6
Instruct FollowingLiveBench
Average Instruction Following Score55.39
6
General EvaluationLiveBench 1125
Score52.1
6
General TasksLiveBench 2024-11-25
Accuracy75.9
5
Mathematical ReasoningLiveBench Math (test)
Score51.95
5
ExaminationLiveBench 2024-11-25
Score70.79
5
Showing 25 of 35 rows