Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Multiple

Benchmarks

Task NameDataset NameSOTA ResultTrend
Video UnderstandingMultiple Aggregate
Average Score69.8
18
Generalist Multi-task EvaluationMultiple (ImageNet-1K, COCO)
Mean Delta-11.8
13
Bayesian uncertainty-aware quantificationMultiple (test)
AE Rank (T=1)1.9
6
Factuality DetectionMultiple TriviaQA, HotpotQA, CSQA
Average AUROC72.9
4
Code GenerationMultiple
Score78.51
3
Controllable Language GenerationMultiple Distributional Constraint
Ctrl0.95
3
Point-estimationMultiple Tabular, Text, Image (test)
Metric-
0
Showing 7 of 7 rows