Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Group Sequence Policy Optimization

About

This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models. Unlike previous algorithms that adopt token-level importance ratios, GSPO defines the importance ratio based on sequence likelihood and performs sequence-level clipping, rewarding, and optimization. We demonstrate that GSPO achieves superior training efficiency and performance compared to the GRPO algorithm, notably stabilizes Mixture-of-Experts (MoE) RL training, and has the potential for simplifying the design of RL infrastructure. These merits of GSPO have contributed to the remarkable improvements in the latest Qwen3 models.

Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, Junyang Lin• 2025

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH500 (test)--
922
Mathematical ReasoningMATH
Accuracy74.2
882
Instruction FollowingAlpacaEval 2.0
Win Rate36.98
752
Mathematical ReasoningMATH 500
Accuracy (Acc)90.3
600
Mathematical ReasoningMATH 500
Accuracy87.6
589
Mathematical ReasoningMATH
Accuracy85.2
535
Mathematical ReasoningAIME 2024
Accuracy73.33
525
Text-to-SQLBIRD (dev)
Execution Accuracy (EA)68.06
477
Mathematical ReasoningMATH 500
Top-1 Accuracy93.7
452
Visual Mathematical ReasoningMathVista
Accuracy81
448
Showing 10 of 417 rows
...

Other info

Follow for update