Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

About

Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced to this field to enhance reasoning accuracy. However, existing CoT-based video reasoning methods primarily rely on text-only information for logical deduction, overlooking critical visual information during the inference process. Inspired by the human cognitive mechanism of reviewing visual segments during inference, we propose VTI-CoT, a Visual-Textual Interleaved CoT framework. VTI-CoT integrates textual reasoning steps with corresponding visual frames. Given the scarcity of visual-textual interleaved CoT in existing datasets, we develop an automated annotation pipeline to construct high-quality multimodal CoT data. Further, reasoning over long-form videos entails increasingly long CoT token sequences, which severely hinders training convergence and efficiency. To address this, we employ Optical Character Recognition (OCR)-based compression techniques to compress CoT supervision signals into a single canvas. Experimental results demonstrate that VTI-CoT achieves state-of-the-art performance among models of the same parameter scale while significantly improving training efficiency.

Shufan Zhang, Ziyue Lin, Bairun Wang, Lei Jin, Xuanding Ding, Xinzhu Ma, Kunlin Yang• 2026

Related benchmarks

TaskDatasetResultRank
Video ReasoningVideo-MME--
73
Video ReasoningMVBench
MVBench Score65.9
56
Video ReasoningTempCompass
Score74.5
46
Extremely long-video understandingLVBench
Score40.5
35
Video ReasoningMMVU mc
Score65.3
32
Long Video ReasoningLongVideoBench
Score55
15
Showing 6 of 6 rows

Other info

Follow for update