Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Evaluation Tasks

Benchmarks

Task NameDataset NameSOTA ResultTrend
Zero-shot EvaluationEvaluation Tasks Zero-shot Aggregate
Avg. Accuracy79.95
74
Language ModelingEvaluation Tasks Zero-shot Average
Zero-shot Average Accuracy60.47
17
Path-Following10,000 Evaluation Tasks Difficult n=2563
Path Length Mean (m)0.32
3
Path-Following10,000 Evaluation Tasks Medium n=4795
Mean Path Length (m)0.57
3
Path-Following10,000 Evaluation Tasks Easy n=2642
Path Length Mean (m)0.53
3
Path-FollowingEvaluation Tasks n=10000 (All)
Mean Path Length (m)0.45
3
Showing 6 of 6 rows