Reinforcement Learning with Parameterized Actions
About
We introduce a model-free algorithm for learning in Markov decision processes with parameterized actions-discrete actions with continuous parameters. At each step the agent must select both which action to use and which parameters to use with that action. We introduce the Q-PAMDP algorithm for learning in these domains, show that it converges to a local optimum, and compare it to direct policy search in the goal-scoring and Platform domains.
Warwick Masson, Pravesh Ranchod, George Konidaris• 2015
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Half Field Offense | Half Field Offense (HFO) (evaluation) | P(Goal)0.00e+0 | 7 | |
| Platform Control | Platform (Evaluation) | Return78.9 | 5 | |
| Robot Soccer | Robot Soccer Goal (Evaluation) | P(Goal)45.2 | 5 |
Showing 3 of 3 rows