Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

WildBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
Creative WritingWildBench
WildBench Score83.9
49
WritingWildBench (test)
Score0.644
32
Instruction FollowingWildBench (test)
Info Seek58.6
27
Open-ended generationWildBench
WildBench0.479
26
Assistant Response GenerationWildBench v2
Win Rate68.4
20
Subjective EvaluationWildBench
Score0.8604
19
General Instruction FollowingWildBench
Score92.6
19
Instruction FollowingWildBench
WB Score63.18
18
Open-ended GenerationWildBench (test)
WildBench Score64.4
17
Creative WritingWildBench (test)
WildBench Score64.4
15
Real-world Query EvaluationWildBench
WildBench Accuracy71.5
14
General ChatWildBench
LLM Judge Score68.16
12
General chatWildBench 2025 (test)
WB-Elo1,062.4
12
Instruction FollowingWildBench 1.0 (test)
WB-Score37.98
8
LLM evaluationWildBench v2
Quality Score64.9
6
Chatbot EvaluationWildBench
Overall Score71.64
6
Open-ended reasoningWildBench
Creative Score57.05
5
Open-ended text generationWildBench
Score-1.7
4
Open-ended instruction followingWildBench v2
Win Rate58.1
3
General Language Model EvaluationWildBench
WildBench Score26.95
2
Showing 20 of 20 rows