Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CommonEval

Benchmarks

Task NameDataset NameSOTA ResultTrend
Audio Question-AnsweringCommonEval
Score91
12
Instruction FollowingCommonEval (ComE)
GPT-score4.08
9
Audio-language generationCommonEval VoiceBench
Score2.94
2
Showing 3 of 3 rows