Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation

About

Recent advances in reinforcement learning (RL) have strengthened the reasoning capabilities of vision-language models (VLMs). However, enhancing policy exploration to better scale test-time compute remains largely underexplored. In addition, VLMs continue to struggle with imperfect visual perception, which in turn affects the subsequent reasoning process. We introduce NoisyRollout, a simple yet effective data augmentation method that addresses these issues by mixing training trajectories from both clean and moderately distorted images. This approach injects perceptual diversity, encouraging better policy exploration and leading to more robust reasoning. A noise annealing schedule gradually reduces distortion strength, aiding exploration early in training while ensuring later stability. Crucially, our method is easy-to-adopt--requiring no additional training cost and no modifications to the RL objective. Extensive experiments on 2 distinct training datasets demonstrate that NoisyRollout achieves state-of-the-art performance among open-source RL-tuned models across 5 out-of-domain reasoning and perception benchmarks. Furthermore, we validate the effectiveness of NoisyRollout across model sizes (7B and 32B), data scales (from 1K to 6K) and image augmentation types (Gaussion noise and rotation), highlighting its generalizability and scalability.

Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du, Longxu Dou, Haonan Wang, Tianyu Pang, Michael Qizhe Shieh• 2025

Related benchmarks

TaskDatasetResultRank
Visual Mathematical ReasoningMathVista
Accuracy78.3
278
Diagram UnderstandingAI2D
Accuracy80.73
247
Mathematical Multimodal ReasoningMathVerse
Accuracy67.8
221
Mathematical Multimodal ReasoningMathVista
Accuracy72.9
218
Visual Mathematical ReasoningMathVision
Accuracy39.82
186
Multimodal Math ReasoningMathVision
Accuracy28.9
183
Multimodal Math ReasoningWeMath
Accuracy71.9
168
Mathematical ReasoningWeMath
Accuracy75.51
161
Mathematical ReasoningMathVision--
144
Multimodal ReasoningMMStar
Accuracy63.93
143
Showing 10 of 57 rows

Other info

Follow for update