Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning
About
Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL (PbRL) offers a promising alternative by learning reward functions from human feedback, but its scalability is hindered by high labeling costs. Inspired by advances in Video Foundation Models (ViFMs), we present Video-based Optimal Transport Preference (VOTP), a semi-supervised framework that learns effective reward functions from only a handful of labels. By leveraging optimal transport to align visual trajectories within the rich representation space of ViFMs, VOTP effectively generates high-fidelity pseudo-labels for large amounts of unlabeled data, substantially reducing human supervision. Extensive experiments across locomotion and manipulation benchmarks demonstrate the superiority of VOTP, which outperforms state-of-the-art offline PbRL methods under limited feedback budgets. We also showcase the robustness of VOTP in the presence of visual distractors and validate its utility on real robotic tasks, where it learns meaningful rewards with minimal human input.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Locomotion | D4RL walker2d-medium-expert v2 | Average Online Return108.1 | 17 | |
| Locomotion | D4RL walker2d medium-replay v2 | Offline Normalized Return66.3 | 16 | |
| Robotic Manipulation | MetaWorld plate-slide v2 | Success Rate57.6 | 11 | |
| Locomotion | D4RL Hopper-medium-expert v2 | Return105.7 | 11 | |
| Robotic Manipulation | MetaWorld door-open v2 | Success Rate84 | 11 | |
| Robotic Manipulation | MetaWorld drawer-open v2 | Success Rate71.2 | 11 | |
| Robotic Manipulation | MetaWorld sweep-into v2 | Success Rate57.6 | 11 | |
| Robotic Manipulation | Sawyer Robot Lift Banana (real-world) | Success Rate80 | 3 | |
| Robotic Manipulation | Sawyer Robot Drawer Open (real-world) | Success Rate70 | 3 |