Multiplayer Nash Preference Optimization

About

Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley-Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences. To address this, recent studies have reframed alignment as a two-player Nash game, giving rise to Nash learning from human feedback (NLHF). While this perspective has inspired algorithms such as INPO, ONPO, and EGPO that offer strong theoretical and empirical guarantees, they remain fundamentally restricted to two-player interactions, introducing a single-opponent bias that fails to capture the full complexity of realistic preference structures. This work introduces Multiplayer Nash Preference Optimization (MNPO), a novel framework that generalizes NLHF to the multiplayer regime. It formulates alignment as an n-player game, where each policy competes against a population of opponents while being regularized toward a reference model. We demonstrate that MNPO inherits the equilibrium guarantees of two-player methods while enabling richer competitive dynamics and improved coverage of diverse preference structures. Comprehensive empirical evaluation shows that MNPO consistently outperforms existing NLHF baselines on instruction-following benchmarks, achieving superior alignment quality under heterogeneous annotator conditions and mixed-policy evaluation scenarios. Together, these results establish MNPO as a principled and scalable framework for aligning LLMs with complex, non-transitive human preferences. Code is available at: https://github.com/smiles724/MNPO

Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi• 2025

Related benchmarks

Task	Dataset	Result
Instruction Following	IFEval	IFEval Accuracy75.26	854
Instruction Following	AlpacaEval 2.0	--	752
Instruction Following	MT-Bench	MT-Bench Score7.52	287
Instruction Following	Arena Hard	Win Rate52.26	263
Knowledge	MMLU	Accuracy75.63	171
Commonsense Reasoning	HellaSwag	HellaSwag Score80.44	62
Commonsense Reasoning	ARC	Accuracy91.23	61
Commonsense Reasoning	TruthfulQA	Accuracy71.8	28
Commonsense Reasoning	WinoGrande	Winogrande Score73.48	22
Mathematical Reasoning	Minerva Math	Last Score49.63	17

Showing 10 of 12 rows

Other info

Follow for update

@wizwand_team Discord