Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

About

Offline-to-online reinforcement learning (RL), by combining the benefits of offline pretraining and online finetuning, promises enhanced sample efficiency and policy performance. However, existing methods, effective as they are, suffer from suboptimal performance, limited adaptability, and unsatisfactory computational efficiency. We propose a novel framework, PROTO, which overcomes the aforementioned limitations by augmenting the standard RL objective with an iteratively evolving regularization term. Performing a trust-region-style update, PROTO yields stable initial finetuning and optimal final performance by gradually evolving the regularization term to relax the constraint strength. By adjusting only a few lines of code, PROTO can bridge any offline policy pretraining and standard off-policy RL finetuning to form a powerful offline-to-online RL pathway, birthing great adaptability to diverse methods. Simple yet elegant, PROTO imposes minimal additional computation and enables highly efficient online finetuning. Extensive experiments demonstrate that PROTO achieves superior performance over SOTA baselines, offering an adaptable and efficient offline-to-online RL framework.

Jianxiong Li, Xiao Hu, Haoran Xu, Jingjing Liu, Xianyuan Zhan, Ya-Qin Zhang• 2023

Related benchmarks

TaskDatasetResultRank
Multi-Agent Reinforcement LearningMAMuJoCo Walker2d 6x1 (test)
Average Episodic Return1.06e+3
13
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 3-Hopper (Medium)
Average Episode Reward1.89e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 6-HalfCheetah (Expert)
Average Episode Reward3.61e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 2-Ant (Medium)
Average Episode Reward1.42e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 6-HalfCheetah (Medium-Expert)
Average Episode Reward3.37e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 6-HalfCheetah (Medium-Replay)
Avg Episode Reward2.33e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo (OMIGA) 6-HalfCheetah Medium
Average Episode Reward2.75e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 2-Ant (Medium-Expert)
Average Episode Reward1.27e+3
11
Multi-Agent Reinforcement LearningMaMuJoCo OMIGA 2-Ant (Medium-Replay)
Average Episode Reward970.4
11
Multi-Agent Reinforcement LearningMA-MuJoCo Ant (4x2) medium
Episode Return1.42e+3
5
Showing 10 of 14 rows

Other info

Follow for update