Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning

About

Reinforcement Learning from Verifiable Rewards (RLVR) has recently become a key paradigm for improving the reasoning abilities of Large Language Models (LLMs), yet it remains limited by sparse binary rewards and its ignorance of model-internal uncertainty. In this paper, we propose ConSteer-RL, a simple yet effective framework that integrates token-level confidence signals derived from model log-probabilities into RLVR training. Specifically, building upon the Group Relative Policy Optimization (GRPO) framework, we construct a confidence-aware reward by aggregating per-token probabilities into a scalar confidence score and incorporating it into an awareness-based reward shaping mechanism that penalizes overconfident errors while reinforcing correct and confident reasoning. Experimental results demonstrate that ConSteer-RL consistently outperforms strong GRPO baselines, achieving average improvements of 2.3%-4.0% across different model scales.

Qing Miao, Yiming Zhao, Jing Yang, Chenxi Liu, Yuehai Chen, Yuewen Liu, Shaoyi Du, Badong Chen• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAIME 2024
Accuracy26.7
394
Mathematical ReasoningOlympiad Bench
Accuracy53.9
254
Mathematical ReasoningMinerva Math
Accuracy49.3
251
Mathematical ReasoningAMC 2023
Accuracy70.9
104
Mathematical ReasoningAIME 2026
AIME 2026 Accuracy23.8
80
Mathematical ReasoningMATH500
Accuracy (%)86.6
56
Showing 6 of 6 rows

Other info

Follow for update