Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

NLU Suite

Benchmarks

Task NameDataset NameSOTA ResultTrend
General Language UnderstandingNLU Suite (MMLU, SST-2, AGNews, 20News, MNLI, SNLI)
Average Accuracy89.9
31
Natural Language UnderstandingNLU suite Zero-Shot (CSQA, SIQA, HS, WG, PIQA, OBQA, ARC:E, ARC:C)
CSQA Accuracy49.8
8
Natural Language InferenceNLU Suite NLI
AUC0.754
4
Showing 3 of 3 rows