Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Distributed Distributional Deterministic Policy Gradients

About

This work adopts the very successful distributional perspective on reinforcement learning and adapts it to the continuous control setting. We combine this within a distributed framework for off-policy learning in order to develop what we call the Distributed Distributional Deep Deterministic Policy Gradient algorithm, D4PG. We also combine this technique with a number of additional, simple improvements such as the use of $N$-step returns and prioritized experience replay. Experimentally we examine the contribution of each of these individual components, and show how they interact, as well as their combined contributions. Our results show that across a wide variety of simple control tasks, difficult manipulation tasks, and a set of hard obstacle-based locomotion tasks the D4PG algorithm achieves state of the art performance.

Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, Timothy Lillicrap• 2018

Related benchmarks

TaskDatasetResultRank
LocomotionDeepMind Control suite Dog-Trot
Final Return146
17
H1sit_hardHumanoid-bench
Total Average Return17
13
H1balance_hardHumanoid-bench
Total Average Return49
13
H1mazeHumanoid-bench
Total Average Return156
13
H1reachHumanoid-bench
Total Average Return2.14e+3
13
Humanoid RunHumanoid-bench
Total Average Return40
13
Walker RunDeepMind Control Suite (DMC)
Average Return636
13
Walker WalkDeepMind Control Suite (DMC)
Total Average Return972
13
Dog-runDeepMind Control Suite (DMC)
Total Average Return106
13
Dog-standDeepMind Control Suite (DMC)
Total Average Return445
13
Showing 10 of 16 rows

Other info

Follow for update