Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

About

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to inform clinical decision-making. We recognize that pruning visual tokens outside the grounding region greatly enhances medical reasoning. However, a united RL framework for active visual token pruning (VTP) and medical multimodal reasoning remains unestablished. Here, we propose a dual-stream RL framework, ViToS, to fulfill token pruning and question answering. ViToS trains one policy model with two task branches, where one focuses on grounding while the other conducts token-sparse reasoning after VTP. Furthermore, we solve the coupled policy learning problem by introducing the cross-feedback sequential optimization, avoiding gradient conflict and facilitating convergence of the shared policy model. Evaluated on seven medical benchmarks, our method reduces visual tokens to 77% of the original sequence length while achieving a 108.27% relative performance on Lingshu-7B and 104.16% relative performance on HuatuoGPT-Vision-7B. Overall, ViToS delivers superior performance and inference speedup, establishing an efficient paradigm for medical multimodal reasoning.

Kaitao Chen, Weiqian Zhao, Jiamin Wu, Qihao Zheng, Shangquan Sun, Chunfeng Song, Xiaosong Wang, Mu Zhou, Mianxin Liu• 2026

Related benchmarks

TaskDatasetResultRank
Medical Visual Question AnsweringSlake
Accuracy90.62
289
Medical Visual Question AnsweringVQA-RAD
Accuracy72.51
251
Medical Visual Question AnsweringPathVQA
Accuracy85.37
103
Multimodal Medical ReasoningVQA-RAD
Accuracy (%)72.51
48
Medical Visual Question AnsweringPMC-VQA
Accuracy63.05
40
Multimodal ReasoningSlake
Accuracy90.62
30
Vision-Language Medical ReasoningPathVQA
Accuracy (%)85.37
30
Medical Visual Question AnsweringMMMU Med
Accuracy78
29
Medical Visual Question AnsweringOmni-Med
Accuracy83.45
18
Medical Visual Question AnsweringMedX-Bench
Accuracy29.8
18
Showing 10 of 17 rows

Other info

Follow for update