Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding

About

Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks.

Jianxiang He, Meisheng Hong, Jungang Li, Weiyu Guo, Xuming Hu, Hui Xiong• 2025

Related benchmarks

TaskDatasetResultRank
Video Question AnsweringLongVideoBench
Accuracy51.2
224
Video UnderstandingLongVideoBench
Accuracy73.89
128
Video Question AnsweringVideo-MME Long
Accuracy55.8
71
Video Question AnsweringVideoMME Medium
Accuracy61.7
53
Video Question AnsweringLONGVIDEOBENCH Medium
Accuracy56.1
24
Keyframe RetrievalLongVideoBench
Precision76.8
7
Downstream Question AnsweringLongVideoBench Medium Text-Related Perception
Accuracy62.07
5
Downstream Question AnsweringLongVideoBench Long Text-Related Perception
Accuracy69.57
5
Keyframe SearchLongVideoBench Text-Related Perception Subsets
Accuracy45
3
Video UnderstandingLONGVIDEOBENCH IMAGE-ONLY
Accuracy71.56
3
Showing 10 of 11 rows

Other info

Follow for update