Hindsight Experience Replay
About
Dealing with sparse rewards is one of the biggest challenges in Reinforcement Learning (RL). We present a novel technique called Hindsight Experience Replay which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering. It can be combined with an arbitrary off-policy RL algorithm and may be seen as a form of implicit curriculum. We demonstrate our approach on the task of manipulating objects with a robotic arm. In particular, we run experiments on three different tasks: pushing, sliding, and pick-and-place, in each case using only binary rewards indicating whether or not the task is completed. Our ablation studies show that Hindsight Experience Replay is a crucial ingredient which makes training possible in these challenging environments. We show that our policies trained on a physics simulation can be deployed on a physical robot and successfully complete the task.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| 2D Interaction Control | Box2D center | Success Rate8.8 | 14 | |
| 2D Interaction Control | Box2D hard velocity | Success Rate15.2 | 14 | |
| Robot Manipulation | MetaWorld pick place | Success Rate0.4 | 14 | |
| Robot Manipulation | MetaWorld push | Success Rate0.4 | 14 | |
| Robot Manipulation | MetaWorld sweep into | Success Rate2 | 14 | |
| 2D Interaction Control | Box2D goal | Success Rate8.6 | 14 | |
| 2D Interaction Control | Box2D hard | Success Rate6.4 | 14 | |
| Robot Manipulation | MetaWorld peg insert | Success Rate0.00e+0 | 14 | |
| 2D Interaction Control | Box2D maze | Success Rate3.1 | 14 | |
| Interaction control | Air Hockey real-transfer | Success Rate12.9 | 14 |