FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

About

Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each question. This static approach limits their ability to adaptively gather visual evidence, leading to suboptimal performance on tasks that require either broad temporal coverage or fine-grained spatial detail. In this paper, we introduce FrameMind, an end-to-end framework trained with reinforcement learning that enables models to dynamically request visual information during reasoning through Frame-Interleaved Chain-of-Thought (FiCOT). Unlike traditional approaches, FrameMind operates in multiple turns where the model alternates between textual reasoning and active visual perception, using tools to extract targeted frames or video clips based on identified knowledge gaps. To train effective dynamic sampling policies, we propose Dynamic Resolution Frame Sampling (DRFS), which exposes models to diverse temporal-spatial trade-offs during learning, and DRFS-GRPO, a group-relative policy optimization algorithm that learns from outcome-based rewards without requiring frame-level annotations. Extensive experiments on challenging benchmarks like MLVU and VideoMME demonstrate that our method significantly outperforms existing models, advancing the state of the art in flexible and efficient video understanding.

Haonan Ge, Yiwei Wang, Kai-Wei Chang, Hang Wu, Yujun Cai• 2025

Related benchmarks

Task	Dataset	Result
Video Understanding	MVBench	Accuracy64.2	635
Video Question Answering	VideoMME	Accuracy60.9	254
Video Question Answering	MLVU	Accuracy48.6	213
Video Understanding	MLVU	Accuracy48.6	147
Video Understanding	Video-MME without subtitles	--	145
Long Video Understanding	Video-MME Long	Accuracy57.5	120
Video Question Answering	MVBench	Accuracy64.2	90
Long Video Understanding	Video-MME Overall	Accuracy60.9	81
Long Video Understanding	MLVU (test)	--	60
Video Understanding	Video-MME w/o sub	Accuracy60.9	46

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord