Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PAWS: Preference Learning with Advantage-Weighted Segments

About

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods typically train utility functions on trajectory or segment-level preferences while relying on per-step utility estimates during policy optimization. This training and inference mismatch induces a distribution shift that severely degrades temporal credit assignment and limits policy learning. We analyze this issue and propose PAWS, a segment-based preference learning method that performs policy updates directly using segment-level advantage functions. By aligning utility training with policy optimization, PAWS preserves trajectory-level preference information and avoids unreliable per-step learning signals. Experiments on simulated robotic manipulation and locomotion tasks demonstrate that PAWS consistently outperforms existing PbRL approaches, highlighting the importance of distribution-consistent preference learning.

Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann• 2026

Related benchmarks

TaskDatasetResultRank
door-openMeta-World
Door Open Success Rate86
28
LocomotionHopper--
20
LocomotionAnt
Reward882
16
LocomotionWalker2D
Reward1.02e+3
16
LocomotionHalfcheetah
Average Episode Return1.48e+3
16
Robotic ManipulationMeta-World modified (test)
Button Press Success84
16
Button pressMeta-World Button Press human preferences
Success Rate57.2
8
Showing 7 of 7 rows

Other info

Follow for update