Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Adaptive Margin RLHF via Preference over Preferences

About

Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinforcement Learning from Human Feedback (RLHF), existing methods typically rely on no margins, fixed margins, or margins that are simplistic functions of preference ratings. However, such formulations often fail to account for the varying strengths of different preferences or they rely on noisy margin information derived from preference ratings. Furthermore, many existing methods that use adaptive margins assume access to accurate preference scores, which can be difficult for humans to provide reliably. We propose leveraging preferences over preferences, that is, annotations indicating which of two preferences reflects a stronger distinction, to infer adaptive margins on a per-datapoint basis. Such preference-over-preference annotations are general and can be incorporated into both standard RLHF reward modeling objectives and direct alignment losses. As a concrete instantiation, we introduce DPO-PoP, an extension to Direct Preference Optimization (DPO) that incorporates adaptive margins from preference-over-preference supervision, enabling improved discriminative and generative performance. Additionally, we show a tradeoff between discriminative and generative performance and propose two sampling strategies for gathering preference-over-preference labels to navigate it.

Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum• 2025

Related benchmarks

TaskDatasetResultRank
Reward ModelingRewardBench
Safety Score81.94
284
Open-ended generationAlpacaEval 2.0
Win Rate13.69
49
LLM EvaluationAlpacaEval 2.0
LC Win Rate14.62
16
Response Preference EvaluationUltraFeedback (test)
Win Rate63
14
Text GenerationUltraRM
Median Adv67.5
5
Showing 5 of 5 rows

Other info

Follow for update