Geometric-Mean Policy Optimization

About

Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unfortunately, GRPO is observed to suffer from unstable policy updates when facing tokens with outlier importance-weighted rewards, which manifest as extreme importance sampling ratios during training. In this study, we propose Geometric-Mean Policy Optimization (GMPO), with the aim to improve the stability of GRPO through suppressing token reward outliers. Instead of optimizing the arithmetic mean, GMPO maximizes the geometric mean of token-level rewards, which is inherently less sensitive to outliers and maintains a more stable range of importance sampling ratio. GMPO is plug-and-play-simply replacing GRPO's arithmetic mean with the geometric mean of token-level rewards, as the latter is inherently less sensitive to outliers. GMPO is theoretically plausible-analysis reveals that both GMPO and GRPO are weighted forms of the policy gradient while the former enjoys more stable weights, which consequently benefits policy optimization and performance. Experiments on multiple mathematical reasoning benchmarks show that GMPO-7B improves the average Pass@1 of GRPO by up to 4.1%, outperforming many state-of-the-art approaches. Code is available at https://github.com/callsys/GMPO.

Yuzhong Zhao, Yue Liu, Junpeng Liu, Jingye Chen, Xun Wu, Yaru Hao, Tengchao Lv, Shaohan Huang, Lei Cui, Qixiang Ye, Fang Wan, Furu Wei• 2025

Related benchmarks

Task	Dataset	Result
Mathematical Reasoning	MATH500 (test)	--	922
Mathematical Reasoning	MATH 500	Top-1 Accuracy91.4	452
Mathematical Reasoning	Minerva	Pass@1 Accuracy37.9	289
Mathematical Reasoning	HMMT 2025	--	241
Mathematical Reasoning	AMC	Accuracy78.3	221
Mathematical Reasoning	Minerva	--	138
Mathematical Reasoning	AMC	Pass@1 Accuracy78.3	119
Mathematical Reasoning	AIME 24	Pass@1 Accuracy46.7	117
Mathematical Reasoning	OlympiadBench	Accuracy0.625	72
Mathematical Reasoning	MATH 500	Pass@176.6	68

Showing 10 of 33 rows

Other info

Follow for update

@wizwand_team Discord