Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GPU-Parallel Multi-Task Reinforcement Learning with Demonstration Guided Policy Optimization

About

Large scale GPU-parallel reinforcement learning has changed what can be trained in robot simulation, yet most systems still optimize one specialist policy per task. We propose a construction methodology for turning structured manipulation task families into GPU-parallel multi-task RL benchmarks, and instantiate it as MT-Libero using LIBERO assets and task predicates in Isaac Lab. The resulting benchmark supports simultaneous reinforcement learning over heterogeneous task suites with parallel rendering, physics randomization, and state-input or visual-input policies. To make such training practical under sparse success signals and limited prior data, we further propose DGPO, an on-policy demonstration guided method that combines importance weighted PPO with adaptive behavior cloning on matched demonstration actions. DGPO enables a tunable preference toward demonstrated task distributions, outperforming both prior-free RL and existing demonstration-based methods while preserving the stability and online improvement benefits of on-policy PPO.

Rui Zhang, Qiwei Wu, Zhengyu Zhang, Tao Li, Yunrong Guo, Junjie Lai, Renjing Xu, Weihua Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Multi-task Robotic ManipulationLIBERO standard suites
Goal Success Rate87.5
10
Multi-task manipulation (simulation only)LIBERO All-40
Throughput3.83e+3
2
End to end PPO training (rollout + training)LIBERO All-40
Throughput2.68e+4
1
End to end PPO training (rollout + training)MT50 rand--
1
Showing 4 of 4 rows

Other info

Follow for update