Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Beyond the Sampled Token: Preserving Candidate Support in RLVR

About

We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction. We formally show that as probability concentrates on the top-$1$ candidate, the expected number of distinct responses collapses to one regardless of the sampling budget $K$. This theoretical implication is further verified by our empirical tracking of top-$N$ candidate probabilities during training, where the top-$1$ candidate progressively dominates while plausible alternatives are suppressed. These findings suggest a key desideratum for effective exploration: \emph{preserving non-negligible probability mass on the top-$N$ candidates}. To this end, we propose Candidate-aware Support Preservation (CaSP), with two complementary designs. Specifically, CaSP redistributes positive gradients among top-$N$ candidates for correct responses, and applies a stronger penalty to the top-$1$ candidate for incorrect responses. Unlike many exploration-oriented methods that improve pass@$K$ at the cost of pass@1, CaSP improves pass@$K$ across the full $K$ spectrum. These gains generalize to 6 math, 2 logical-reasoning, and 2 coding benchmarks, and scales to 32B-parameter models and sampling budgets up to $K=1024$, positioning it as a principled, candidate-level approach for RLVR exploration.

Ruotian Peng, Yi Ren, Zhouliang Yu, Weiyang Liu, Yandong Wen• 2025

Related benchmarks

TaskDatasetResultRank
Code GenerationHumanEval
Pass@156.2
171
Mathematical ReasoningAIME25 (test)
Pass@127.5
45
Mathematical ReasoningAIME 2024
Pass@132.8
22
Mathematical ReasoningAIME 2025
Pass@117.8
22
Mathematical ReasoningAMC 2023
Pass@168.5
22
Mathematical ReasoningMATH 500
Pass@183.8
22
Mathematical ReasoningMinerva Math
Pass@145.2
22
Mathematical ReasoningOlympiadBench
Pass@147.1
22
Mathematical ReasoningOmniMath (test)--
20
Code GenerationLiveCodeBench
Rate @32 Score46.5
17
Showing 10 of 18 rows

Other info

Follow for update