Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Cleanup

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multi-agent policy synthesisCleanup
U Score2.75
9
Multi-agent Social Dilemma Equality EvaluationCleanup
Equality Score (E)95.9
9
Social Outcome EvaluationCleanup Normal (train test)
Outcome U0.6
4
Social Outcome EvaluationCleanup Hard (train test)
Outcome U31
4
Robot Plan ExecutionCleanup real-world
Success Rate (New Objects)3
2
Showing 5 of 5 rows