Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Hindsight Experience Replay

About

Dealing with sparse rewards is one of the biggest challenges in Reinforcement Learning (RL). We present a novel technique called Hindsight Experience Replay which allows sample-efficient learning from rewards which are sparse and binary and therefore avoid the need for complicated reward engineering. It can be combined with an arbitrary off-policy RL algorithm and may be seen as a form of implicit curriculum. We demonstrate our approach on the task of manipulating objects with a robotic arm. In particular, we run experiments on three different tasks: pushing, sliding, and pick-and-place, in each case using only binary rewards indicating whether or not the task is completed. Our ablation studies show that Hindsight Experience Replay is a crucial ingredient which makes training possible in these challenging environments. We show that our policies trained on a physics simulation can be deployed on a physical robot and successfully complete the task.

Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, Wojciech Zaremba• 2017

Related benchmarks

TaskDatasetResultRank
2D Interaction ControlBox2D center
Success Rate8.8
14
2D Interaction ControlBox2D hard velocity
Success Rate15.2
14
Robot ManipulationMetaWorld pick place
Success Rate0.4
14
Robot ManipulationMetaWorld push
Success Rate0.4
14
Robot ManipulationMetaWorld sweep into
Success Rate2
14
2D Interaction ControlBox2D goal
Success Rate8.6
14
2D Interaction ControlBox2D hard
Success Rate6.4
14
Robot ManipulationMetaWorld peg insert
Success Rate0.00e+0
14
2D Interaction ControlBox2D maze
Success Rate3.1
14
Interaction controlAir Hockey real-transfer
Success Rate12.9
14
Showing 10 of 52 rows

Other info

Follow for update