Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Stick-Breaking Policy Learning in Dec-POMDPs

About

Expectation maximization (EM) has recently been shown to be an efficient algorithm for learning finite-state controllers (FSCs) in large decentralized POMDPs (Dec-POMDPs). However, current methods use fixed-size FSCs and often converge to maxima that are far from optimal. This paper considers a variable-size FSC to represent the local policy of each agent. These variable-size FSCs are constructed using a stick-breaking prior, leading to a new framework called \emph{decentralized stick-breaking policy representation} (Dec-SBPR). This approach learns the controller parameters with a variational Bayesian algorithm without having to assume that the Dec-POMDP model is available. The performance of Dec-SBPR is demonstrated on several benchmark problems, showing that the algorithm scales to large problems while outperforming other state-of-the-art methods.

Miao Liu, Christopher Amato, Xuejun Liao, Lawrence Carin, Jonathan P. How• 2015

Related benchmarks

TaskDatasetResultRank
Dec-POMDP Policy Learning and PlanningMARS ROVERS 256, 6, 8
Policy Value20.62
5
Dec-POMDP Policy Learning and PlanningDEC-TIGER 2, 3, 3
Policy Value-18.63
5
Dec-POMDP Policy Learning and PlanningRECYCLING ROBOTS (3, 3, 2)
Policy Value31.26
5
Dec-POMDP Policy Learning and PlanningBOX PUSHING 100, 4, 5
Policy Value77.65
5
Dec-POMDP Policy Learning and PlanningBROADCAST (4, 2, 5)
Policy Value9.27
4
Showing 5 of 5 rows

Other info

Follow for update