Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

About

Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff between computational efficiency and sample efficiency in value estimation and policy learning. We introduce BASIS, a critic-free post-training algorithm designed to address this tradeoff. At each online training step, BASIS samples only one rollout per prompt, but leverages rich information across prompts in the entire batch to improve value function estimation. Our experiments demonstrate that BASIS reduces MSE in value function estimation by 69% compared to REINFORCE++, a representative single-rollout baseline, and achieves lower MSE with one rollout than group mean estimators with 8 rollouts. This improvement in value estimation translates to better policy optimization: using substantially less training time, BASIS achieves performance close to multi-rollout GRPO-type baselines and often outperforms single-rollout REINFORCE-type baselines.

Shijin Gong, Erhan Xu, Kai Ye, Francesco Quinzan, Giulia Livieri, Chengchun Shi• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAIME 2024
Accuracy33.8
525
Mathematical ReasoningAIME 2025
Accuracy28.4
353
Mathematical ReasoningOlympiad Bench
Accuracy57.4
254
Mathematical ReasoningHMMT 2025
Accuracy11.4
241
Mathematical ReasoningMATH 500
Accuracy89.2
183
Mathematical ReasoningMinerva Math
Accuracy46
124
Mathematical ReasoningMath Benchmarks Average
Accuracy (ACC)48.3
47
Showing 7 of 7 rows

Other info

Follow for update