Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation

About

Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workflows often lead to low hardware utilization and slow training on existing systems. In this paper, we present RLinf, a high-performance RL training system based on our key observation that the major roadblock to efficient RL training lies in system flexibility. To maximize flexibility and efficiency, RLinf is built atop a novel RL system design paradigm called macro-to-micro flow transformation (M2Flow), which automatically breaks down high-level, easy-to-compose RL workflows at both the temporal and spatial dimensions, and recomposes them into optimized execution flows. Supported by RLinf worker's adaptive communication capability, we devise context switching and elastic pipelining to realize M2Flow transformation, and a profiling-guided scheduling policy to generate optimal execution plans. Extensive evaluations on both reasoning RL and embodied RL tasks demonstrate that RLinf consistently outperforms state-of-the-art systems, achieving $1.07\times-2.43\times$ speedup in end-to-end training throughput.

Chao Yu, Yuanqing Wang, Zhen Guo, Hao Lin, Si Xu, Hongzhi Zang, Quanlu Zhang, Yongji Wu, Chunyang Zhu, Junhao Hu, Zixiao Huang, Mingjie Wei, Yuqing Xie, Ke Yang, Bo Dai, Zhexuan Xu, Jiakun Du, Xiangyuan Wang, Xu Fu, Letong Shi, Zhihao Liu, Kang Chen, Weilin Liu, Gang Liu, Boxun Li, Jianlei Yang, Zhi Yang, Guohao Dai, Yu Wang• 2025

Related benchmarks

TaskDatasetResultRank
VLA Training Throughput (pi_0)ManiSkill
Training Throughput370.3
15
VLA Training Throughput (GR00T N1.5)LIBERO
Throughput (samples/s)1.13e+3
15
VLA Training Throughput (pi_0.5)LIBERO
Training Throughput (pi_0.5)703.9
15
Robot ManipulationLIBERO 4Tasks
Final Success Rate93
4
Robot ManipulationALOHA-Cube
Final Success Rate91
4
Organizing toolboxFetch Simulation Organizing toolbox
Success Rate0.00e+0
3
Soda-can disposalFetch Simulation Soda-can disposal
Success Rate35
3
Sorting vegetablesFetch Simulation Sorting vegetables
Success Rate0.00e+0
3
Showing 8 of 8 rows

Other info

Follow for update