Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

A Regret Minimization Framework on Preference Learning in Large Language Models

About

Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals. However, many realistic language tasks are difficult to equip with reliable verifiers, motivating a growing reliance on reinforcement learning from human feedback (RLHF). In this setting, we argue that a closer examination of how human feedback should be interpreted is essential. We introduce Regret-based Preference Optimization $(\textbf{RePO})$, which reframes RLHF through $\textit{regret minimization}$ rather than reward maximization. Human preferences are often shaped by $\textit{prospective}$ anticipation of outcomes and $\textit{counterfactual}$ comparisons to alternative behaviors, rather than by immediate, outcome-independent utility. $\textbf{RePO}$ captures this structure by modeling preferences as behavior-conditioned assessments of relative suboptimality. Experiments on mathematical reasoning benchmarks and human preference datasets demonstrate consistent performance gains, indicating that $\textbf{RePO}$ is an effective and human-aligned approach for training large language models.

Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim, Yu Jin Kim, Youngsoo Jang, Moontae Lee, Jungwoo Lee• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAMC 23
Accuracy22.5
83
Mathematical ReasoningGSM8K
Accuracy91.05
80
Preference EvaluationAlpacaEval 2
WR (%)60.12
64
Mathematical ReasoningAMC23
Mean Accuracy45
42
Human Preference AlignmentMT-Bench--
20
Mathematical ReasoningGSM8K
Pass@1 Accuracy89.49
19
Human Preference AlignmentArena Hard
Win Rate (%)60.1
16
Mathematical ReasoningMATH
Accuracy (Mean)63.95
14
Mathematical ReasoningMATH 500
Mean Accuracy (MATH 500)63.68
14
Mathematical ReasoningMinerva
Mean Accuracy21.69
14
Showing 10 of 10 rows

Other info

Follow for update