Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Batch Policy Learning under Constraints

About

When learning policies for real-world domains, two important questions arise: (i) how to efficiently use pre-collected off-policy, non-optimal behavior data; and (ii) how to mediate among different competing objectives and constraints. We thus study the problem of batch policy learning under multiple constraints, and offer a systematic solution. We first propose a flexible meta-algorithm that admits any batch reinforcement learning and online learning procedure as subroutines. We then present a specific algorithmic instantiation and provide performance guarantees for the main objective and all constraints. To certify constraint satisfaction, we propose a new and simple method for off-policy policy evaluation (OPE) and derive PAC-style bounds. Our algorithm achieves strong empirical results in different domains, including in a challenging problem of simulated car driving subject to multiple constraints such as lane keeping and smooth driving. We also show experimentally that our OPE method outperforms other popular OPE techniques on a standalone basis, especially in a high-dimensional setting.

Hoang M. Le, Cameron Voloshin, Yisong Yue• 2019

Related benchmarks

TaskDatasetResultRank
Offline Policy SelectionSepsis Simulator simulated (test)
AE0.07
18
Policy RankingPendulum
Regret0.02
8
Policy RankingAnt
Regret0.00e+0
8
Policy RankingHopper
Regret0.54
8
Policy Rankingcheetah
Regret1.29
8
Offline Policy EvaluationSepsis-MDP (test)
MSE0.2448
8
Policy RankingLunarLander
Regret2.51
8
Policy RankingWalker
Regret1.18
8
Offline Policy EvaluationSepsis-POMDP (test)
MSE0.3931
8
Off-policy EvaluationALFWorld (iter1)
Rho0.82
6
Showing 10 of 18 rows

Other info

Follow for update