Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

FrameOracle: Learning What to See and How Much to See in Videos

About

Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as cross-modal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16-frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding.

Chaoyu Li, Tianzhi Li, Fei Tao, Zhenyu Zhao, Ziqian Wu, Maozheng Zhao, Juntong Song, Cheng Niu, Pooyan Fazli• 2025

Related benchmarks

TaskDatasetResultRank
Long Video UnderstandingLongVideoBench (val)
Accuracy65.2
282
Long Video UnderstandingMLVU
Accuracy66.3
265
Video Question AnsweringEgoSchema
Accuracy62.8
194
Video UnderstandingLVB
Accuracy56.5
101
Long Video UnderstandingVideoMME
Accuracy58.1
97
Video UnderstandingVideo-MME
Accuracy58.9
90
Long-form Video UnderstandingEgoSchema
Accuracy63.4
69
Video UnderstandingLongVideoBench--
59
Multi-modal Video EvaluationVideo-MME
Accuracy69.1
57
Egocentric Video UnderstandingEgoSchema--
42
Showing 10 of 17 rows

Other info

Follow for update