Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

About

Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce variance, it lacks a theoretical explanation of why it works and whether it is important or potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple continuations are sampled for each thought. Using the multivariate delta method, we reveal a sampling-dimension asymmetry. Increasing sampled thoughts ($K$) leaves a strictly positive estimation-variance floor, whereas increasing continuations per thought ($M$) drives the leading-order estimation variance to zero at rate $1/M$. This implies that, within the fixed-temperature GRPO-style estimator without value models studied here, accurate thought-level advantage estimation cannot be achieved by scaling thought sampling alone, making continuation-level branching a principled and potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for its effectiveness and potential necessity, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across vision domains and under different model architectures and sizes.

Hongcheng Wang, Yinuo Huang, Sukai Wang, Guanghui Ren, Hao Dong• 2025

Related benchmarks

TaskDatasetResultRank
Math ReasoningGSM8K
Pass@1 Accuracy70.58
61
Trajectory PredictionTrajectory Prediction
DFD151.1
15
OCR-based VQADocVQA
ANLS94.22
8
Math generationAIME
Pass@1014.7
8
Code GenerationLiveBench
Pass@1011.69
8
OCR-based VQAInfographicsVQA
ANLS76.69
8
OCR-based VQAST-VQA
ANLS72.48
8
General Question AnsweringExpertQA
Reward0.214
8
Manipulating Point PredictionManipulating Point Prediction (Seen)
Success Rate31.4
7
Manipulating Point PredictionManipulating Point Prediction (Unseen)
Success Rate16
7
Showing 10 of 12 rows

Other info

Follow for update