Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

InvertedDoublePendulum

Benchmarks

Task NameDataset NameSOTA ResultTrend
Reinforcement LearningInvertedDoublePendulum v5
Avg AUC (z-scored)1.29
13
Reinforcement LearningInvertedDoublePendulum v3
Average Final Return9,360
7
Continuous ControlInvertedDoublePendulum v1 (train)
Max Average Return9,355.52
7
Reinforcement LearningInvertedDoublePendulum Gymnasium
Mean Best Reward3,609.37
5
Reinforcement Learning Surrogate ModelingInvertedDoublePendulum (IDP) (test)
Reward Ratio (%)103
4
Policy ImprovementInvertedDoublePendulum (IDP)
Success Rate38
4
Reinforcement LearningInvertedDoublePendulum v4
Average Episodic Reward9,167.5
4
Continuous ControlInvertedDoublePendulum v5
Average Episodic Reward9,349.2
2
Showing 8 of 8 rows