Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregation

About

While reinforcement learning (RL) significantly enhances LLM reasoning, its efficacy is severely undermined by Pre-RL data overlap, where RL datasets overlap with pretraining or SFT corpora, causing models to exploit shortcuts by memorizing correct answers and fabricating post-hoc reasoning. To address this, we introduce HIPPO, a novel RL framework that integrates hint-injected aggregation with a tailored pairwise reward model. By utilizing hint injection to deliberately trigger overlap-induced behaviors, the resulting traces naturally serve as explicit anchors for pairwise comparison. This provides highly discriminable preference signals, enabling a lightweight judge model to reliably distinguish genuine reasoning deduction from shortcut-driven rationalization, while the pairwise formulation ensures stable and robust optimization compared to standard PRMs. Extensive experiments demonstrate that HIPPO yields substantial improvements over standard baselines and generalizes effectively to out-of-distribution general tasks, showing it extracts authentic, transferable reasoning skills rather than superficial shortcut patterns.

Jiuheng Lin, Chen Zhang, Yansong Feng• 2026

Related benchmarks

TaskDatasetResultRank
Medical Question AnsweringMedMCQA
Accuracy53.3
591
Medical ReasoningMedMCQA
Accuracy57
72
Mathematical ReasoningTheoremQA
Accuracy51.2
67
Medical ReasoningMedQA
Accuracy61.7
61
Mathematical ReasoningDeepScaleR
Accuracy62.1
34
Mathematical ReasoningMATH 500
Accuracy (MATH-500)74.6
33
Multi-task Language UnderstandingMMLU-Pro
Biology Score75
23
General ReasoningMMLU-Pro
Accuracy58.8
22
Mathematical ReasoningMATH 500
Accuracy69.3
18
Mathematical ReasoningCARP-EN
Average@20.593
17
Showing 10 of 13 rows

Other info

Follow for update