Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Reinforcement Learning with Parameterized Actions

About

We introduce a model-free algorithm for learning in Markov decision processes with parameterized actions-discrete actions with continuous parameters. At each step the agent must select both which action to use and which parameters to use with that action. We introduce the Q-PAMDP algorithm for learning in these domains, show that it converges to a local optimum, and compare it to direct policy search in the goal-scoring and Platform domains.

Warwick Masson, Pravesh Ranchod, George Konidaris• 2015

Related benchmarks

TaskDatasetResultRank
Half Field OffenseHalf Field Offense (HFO) (evaluation)
P(Goal)0.00e+0
7
Platform ControlPlatform (Evaluation)
Return78.9
5
Robot SoccerRobot Soccer Goal (Evaluation)
P(Goal)45.2
5
Showing 3 of 3 rows

Other info

Follow for update