Maximum a Posteriori Policy Optimisation
About
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show that several existing methods can directly be related to our derivation. We develop two off-policy algorithms and demonstrate that they are competitive with the state-of-the-art in deep reinforcement learning. In particular, for continuous control, our method outperforms existing methods with respect to sample efficiency, premature convergence and robustness to hyperparameter settings while achieving similar or better final performance.
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, Martin Riedmiller• 2018
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Locomotion | DeepMind Control suite Dog-Trot | Final Return98 | 17 | |
| H1balance_hard | Humanoid-bench | Total Average Return49 | 13 | |
| H1maze | Humanoid-bench | Total Average Return145 | 13 | |
| H1reach | Humanoid-bench | Total Average Return2.02e+3 | 13 | |
| Humanoid-stand | Humanoid-bench | Total Return9 | 13 | |
| Dog-run | DeepMind Control Suite (DMC) | Total Average Return83 | 13 | |
| Dog-stand | DeepMind Control Suite (DMC) | Total Average Return390 | 13 | |
| H1crawl | Humanoid-bench | Total Average Return212 | 13 | |
| H1sit_hard | Humanoid-bench | Total Average Return11 | 13 | |
| Humanoid Run | Humanoid-bench | Total Average Return2 | 13 |
Showing 10 of 15 rows