Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Policy-labeled Preference Learning: Is Preference Enough for RLHF?

About

To design rewards that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing policies via reinforcement learning algorithms. However, existing RLHF methods often misinterpret trajectories as being generated by an optimal policy, causing inaccurate likelihood estimation and suboptimal learning. Inspired by Direct Preference Optimization framework which directly learns optimal policy without explicit reward, we propose policy-labeled preference learning (PPL), to resolve likelihood mismatch issues by modeling human preferences with regret, which reflects behavior policy information. We also provide a contrastive KL regularization, derived from regret-based principles, to enhance RLHF in sequential decision making. Experiments in high-dimensional continuous control tasks demonstrate PPL's significant improvements in offline RLHF performance and its effectiveness in online settings.

Taehyun Cho, Seokhun Ju, Seungyub Han, Dohyeong Kim, Kyungjae Lee, Jungwoo Lee• 2025

Related benchmarks

TaskDatasetResultRank
door-openMeta-World
Door Open Success Rate54.8
28
LocomotionHopper--
20
Robotic ManipulationMeta-World modified (test)
Button Press Success86
16
LocomotionWalker2D
Reward476
16
LocomotionAnt
Reward581
16
LocomotionHalfcheetah
Average Episode Return998
16
Button pressMeta-World Button Press human preferences
Success Rate49
8
Showing 7 of 7 rows

Other info

Follow for update