Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Distributionally Robust Listwise Preference Optimization

About

Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from $K!$ enumeration to $O(K\log K)$. This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global $\epsilon$-suboptimality with $O(\epsilon^{-2})$ sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and $\widetilde O(\epsilon^{-2})$ Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model and external GPT-4 judge metrics.

Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen• 2026

Related benchmarks

TaskDatasetResultRank
RankingUltraFeedback clean (held-out)
Kendall’s Tau0.356
60
Preference AlignmentU10 (held-out evaluation set)
Delta Reward610.7
15
Reward ModelingRewardBench External Evaluation
Chat Score93
6
Reward ModelingUltraFeedback Top-rank 0.4 noisy (test)
Kendall's Tau0.356
4
Reward ModelingUltraFeedback Near-tie 0.4 noisy (test)
Kendall’s Tau0.37
4
Reward ModelingUltraFeedback Clean (test)
Kendall's τ0.38
4
Showing 6 of 6 rows

Other info

Follow for update