Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Learning to Refine Hidden States for Reliable LLM Reasoning

About

Large language models show strong reasoning ability, but their internal reasoning process can remain unstable in complex multi-step settings, where early hidden-state errors may propagate to incorrect predictions. We propose ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations before decoding. ReLAR maintains a compact latent reasoning state and uses learned depth and action controllers to adaptively determine both the number and direction of refinement steps. The controllers are trained with a policy gradient objective based on step-wise likelihood improvement, enabling efficient input-dependent reasoning without explicit chain-of-thought generation. Experiments on medical, mathematical, multi-hop reasoning, and open-ended generation benchmarks show that ReLAR improves accuracy, generation quality, and reasoning stability with substantially lower inference overhead than explicit reasoning baselines.

Chia-Hsuan Hsu, Jui-Ming Yao• 2026

Related benchmarks

TaskDatasetResultRank
Medical ReasoningPubMedQA
Accuracy79.23
48
Mathematical ReasoningGSM-Hard
Accuracy (GSM-Hard)48.57
24
Multi-hop ReasoningHotpotQA
Accuracy59.64
24
Mathematical ReasoningGSM8K
Accuracy71.28
24
Text GenerationCommonGen
ROUGE-L38.92
17
Open-ended generationWritingPrompts
BERTScore87.8
8
Showing 6 of 6 rows

Other info

Follow for update