Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MultiChallenge

Benchmarks

Task NameDataset NameSOTA ResultTrend
Instruction FollowingMultiChallenge
Score69.3
26
Instruction FollowingMultiChallenge (Out-of-Domain)
Overall Score38.5
23
Reverse Chain-of-Thought GenerationMultiChallenge
Score45
20
General-purpose BehaviorMultiChallenge
Score58.6
7
Turn-level correlation with human ratingsMultiChallenge
Spearman Correlation0.74
6
Session-level correlation with human ratingsMultiChallenge
Spearman Correlation (ρ)0.74
6
Long-context & Multi-turn DialogueMultichallenge
Score39.71
4
Multi-turn Dialogue ReasoningMultiChallenge
Accuracy32.97
4
Medical Instruction FollowingMultiChallenge
Pass@166.8
4
Relational UnderstandingMultiChallenge
IM Score20.35
2
Showing 10 of 10 rows