Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

What are Key Factors for Updates in RL for LLM Reasoning?

About

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existing work is guided by heuristic intuition, leading to divergent algorithmic choices, even contradictory ones that nevertheless report empirical gains. To better understand this phenomenon, we conduct a theoretical analysis of RLVR updates. Our study reveals that differences in off-policy degree, determined by the number of gradient steps per rollout, substantially affect the distribution of importance sampling ratios and their clipping behavior, thereby altering which tokens dominate the update. Building on this insight, we characterize gradient expectation as the central quantity governing update dynamics and analyze the roles of token probability, advantage, and importance sampling ratio. Motivated by these findings, we propose Adaptive Clip Policy Optimization (ACPO), which adjusts clipping boundaries across token groups according to the empirical variance of their importance sampling ratios. Experiments on 3B and 7B models across diverse reasoning benchmarks, spanning mathematical problem solving, tabular QA, and logic puzzles, demonstrate that ACPO outperforms strong baselines such as DAPO and CISPO. These results demonstrate that principled, analysis-driven approaches yield more robust and effective RLVR methods. Code is available in: https://github.com/Control-derek/ACPO

Peidong Wang, Demi Wang, Xufang Luo, Jiahang Xu, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical Problem SolvingMATH500
Accuracy80.48
96
Mathematical Problem SolvingAIME 25
Accuracy14.37
84
Math problem solvingOlympiadBench
Accuracy42.83
50
Math problem solvingAIME 24
Accuracy16.77
37
Tabular Question AnsweringHiTab (val)
Validation Reward (mean@4)69.83
32
Mathematical Problem SolvingAMC 2023
Accuracy (avg@k)54.37
27
Arithmetic ReasoningCountdown (val)
Validation Reward (mean@4)76.27
13
Mathematical Problem SolvingMinerva
Accuracy31.66
13
Mathematical Problem SolvingORZ-57K AIME24/25, Minerva, Math500, AMC, OlympiadBench (test)
Accuracy (Minerva)31.66
3
Showing 9 of 9 rows

Other info

Follow for update