Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Maximum a Posteriori Policy Optimisation

About

We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are competitive with the state-of-the-art in deep reinforcement learning. In particular, for continuous control, our method outperforms existing methods with respect to sample efficiency, premature convergence and robustness to hyperparameter settings while achieving similar or better final performance.

Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, Martin Riedmiller• 2018

Related benchmarks

TaskDatasetResultRank
LocomotionDeepMind Control suite Dog-Trot
Final Return98
17
H1balance_hardHumanoid-bench
Total Average Return49
13
H1mazeHumanoid-bench
Total Average Return145
13
H1reachHumanoid-bench
Total Average Return2.02e+3
13
Humanoid-standHumanoid-bench
Total Return9
13
Dog-runDeepMind Control Suite (DMC)
Total Average Return83
13
Dog-standDeepMind Control Suite (DMC)
Total Average Return390
13
H1crawlHumanoid-bench
Total Average Return212
13
H1sit_hardHumanoid-bench
Total Average Return11
13
Humanoid RunHumanoid-bench
Total Average Return2
13
Showing 10 of 15 rows

Other info

Follow for update