RLPR: Extrapolating RLVR to General Domains without Verifiers

About

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates promising potential in advancing the reasoning capabilities of LLMs. However, its success remains largely confined to mathematical and code domains. This primary limitation stems from the heavy reliance on domain-specific verifiers, which results in prohibitive complexity and limited scalability. To address the challenge, our key observation is that LLM's intrinsic probability of generating a correct free-form answer directly indicates its own evaluation of the reasoning reward (i.e., how well the reasoning process leads to the correct answer). Building on this insight, we propose RLPR, a simple verifier-free framework that extrapolates RLVR to broader general domains. RLPR uses the LLM's own token probability scores for reference answers as the reward signal and maximizes the expected reward during training. We find that addressing the high variance of this noisy probability reward is crucial to make it work, and propose prob-to-reward and stabilizing methods to ensure a precise and stable reward from LLM intrinsic probabilities. Comprehensive experiments in four general-domain benchmarks and three mathematical benchmarks show that RLPR consistently improves reasoning capabilities in both areas for Gemma, Llama, and Qwen based models. Notably, RLPR outperforms concurrent VeriFree by 7.6 points on TheoremQA and 7.5 points on Minerva, and even surpasses strong verifier-model-dependent approaches General-Reasoner by 1.6 average points across seven benchmarks.

Tianyu Yu, Bo Ji, Shouli Wang, Shu Yao, Zefan Wang, Ganqu Cui, Lifan Yuan, Ning Ding, Yuan Yao, Zhiyuan Liu, Maosong Sun, Tat-Seng Chua• 2025

Related benchmarks

Task	Dataset	Result
Instruction Following	IFEval	--	836
Instruction Following	AlpacaEval 2.0	--	722
Mathematical Reasoning	AIME 2024	Accuracy24.5	370
General Knowledge	MMLU	MMLU General Knowledge Accuracy68.7	307
Mathematical Problem Solving	MATH	Accuracy52.7	229
Mathematical Reasoning	AMC	Accuracy53	221
Mathematical Reasoning	AIME 2025	Accuracy30.78	214
General Reasoning	MMLU-Pro	Accuracy45.7	201
Knowledge Reasoning	MMLU-Pro	--	120
Code	HumanEval	HumanEval Accuracy72.3	118

Showing 10 of 53 rows

Other info

Follow for update

@wizwand_team Discord