Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

About

Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning. To address this problem, we propose VOLD, a framework to transfer reasoning capabilities from text-only teacher models to VLM student models. To this end, VOLD combines reinforcement learning via Group Relative Policy Optimization (GRPO) with on-policy distillation, which allows the student reasoning traces to be guided by the teacher model, resulting in a significant gain over using GRPO alone. We further show that a cold-start alignment is essential for an effective transfer during the online training phase in this scenario and that without sufficient distributional alignment between teacher and student, on-policy distillation fails to provide meaningful guidance. We evaluate VOLD across diverse benchmarks including MMMU-Pro, MathVision, MathVista, and LogicVista, showing that VOLD outperforms the baseline model significantly and improves over the state of the art by a margin. Our ablation shows the importance of a cold-start alignment via SFT for on-policy distillation with a text-only teacher.

Walid Bousselham, Hilde Kuehne, Cordelia Schmid• 2025

Related benchmarks

TaskDatasetResultRank
Multimodal Mathematical ReasoningDynaMath
Accuracy (DynaMath)50.7
45
Multimodal Mathematical ReasoningMathVerse 1.0 (test)
Score37.9
17
Multimodal Logical ReasoningLogicVista v1.0 (test)
Accuracy45
5
Multimodal Mathematical ReasoningMathVision v1.0 (test)
Accuracy28
5
Multimodal Mathematical ReasoningWeMath v1.0 (test)
Accuracy31.8
5
Multimodal ReasoningMMMU-Pro Vision v1.0 (test)
Accuracy32
5
Multimodal Mathematical ReasoningMathVista v1.0 (test)
Accuracy61.9
5
Multimodal ReasoningMMStar v1.0 (test)
Accuracy55.2
5
Showing 8 of 8 rows

Other info

Follow for update