Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ReFPO: Reflow Regularization for Flow Matching Policy Gradients

About

We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We uncover a key structural property: the gradient updates in Flow Matching Policy Gradients (FPO) can be interpreted as an implicit advantage-weighted Reflow process, providing a new geometric perspective on flow-based policy gradients. Building on this insight, ReFPO introduces an explicit geometric regularizer that can be implemented with a single line of code change without incurring additional computational overhead or auxiliary distillation stages. By synergizing advantage-guided updates with path rectification, our method reduces CFM proxy-ratio spikes, stabilizes PPO-style training, and enables high-fidelity one-step inference that often matches or exceeds multi-step performance. We experimentally demonstrate that ReFPO improves average performance and discretization robustness across GridWorld, MuJoCo Playground, and high-dimensional Humanoid Control tasks, providing a scalable and stable approach for generative policies in complex physical simulations.

Ge Wang, Yibo Peng, Fan Feng, Shenhao Yan, Chengsi Yao, Jiahao Yang, Honghao Cai, Yiming Zhao, Xi Li, Jinke Ren, Shuguang Cui, Yatong Han, Zhen Li• 2026

Related benchmarks

TaskDatasetResultRank
Continuous ControlMuJoCo Playground 10 tasks
10-step Reward686
11
Humanoid ControlHumanoid Control Root + Hands (test)
Success Rate74.4
5
Humanoid ControlHumanoid Control Root (test)
Success Rate58.7
5
Humanoid ControlHumanoid Control All joints (test)
Success Rate97.3
5
Showing 4 of 4 rows

Other info

Follow for update