PyVision-RL: Forging Open Agentic Vision Models via RL

About

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. Our approach combines an oversampling-filtering-ranking rollout strategy with an accumulative tool reward to prevent collapse and encourage multi-turn tool use. Using a unified training pipeline, we develop PyVision-Image and PyVision-Video for image and video understanding. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames during reasoning to significantly reduce visual token usage. Experiments show strong performance and improved efficiency, demonstrating that sustained interaction and on-demand visual processing are critical for scalable multimodal agents.

Shitian Zhao, Shaoheng Lin, Ming Li, Haoquan Zhang, Wenshuo Peng, Kaipeng Zhang, Chen Wei• 2026

Related benchmarks

Task	Dataset	Result
Multimodal Reasoning	WeMath	Accuracy47.7	171
Multimodal Reasoning	MathVision	Accuracy28.7	162
Multimodal Reasoning	MathVerse	Accuracy55.8	130
High-resolution Visual Understanding	HR-Bench-8K	--	83
Multimodal Reasoning	DynaMath	Accuracy61.6	72
Document Visual Question Answering	DocVQA v1.0 (test)	--	49
Spatial Reasoning	VSI-Bench (test)	Avg Score44	31
Visual Search	HR-Bench-4K	Accuracy78.1	29
Visual Search	HR-Bench-8K	Accuracy74.3	29
Tool Use	VerlTool IID Tools	Att.65	11

Showing 10 of 16 rows

Other info

GitHub

Follow for update

@wizwand_team Discord