HoliTom: Holistic Token Merging for Fast Video Large Language Models

About

Video large language models (video LLMs) excel at video comprehension but face significant computational inefficiency due to redundant video tokens. Existing token pruning methods offer solutions. However, approaches operating within the LLM (inner-LLM pruning), such as FastV, incur intrinsic computational overhead in shallow layers. In contrast, methods performing token pruning before the LLM (outer-LLM pruning) primarily address spatial redundancy within individual frames or limited temporal windows, neglecting the crucial global temporal dynamics and correlations across longer video sequences. This leads to sub-optimal spatio-temporal reduction and does not leverage video compressibility fully. Crucially, the synergistic potential and mutual influence of combining these strategies remain unexplored. To further reduce redundancy, we introduce HoliTom, a novel training-free holistic token merging framework. HoliTom employs outer-LLM pruning through global redundancy-aware temporal segmentation, followed by spatial-temporal merging to reduce visual tokens by over 90%, significantly alleviating the LLM's computational burden. Complementing this, we introduce a robust inner-LLM token similarity-based merging approach, designed for superior performance and compatibility with outer-LLM pruning. Evaluations demonstrate our method's promising efficiency-performance trade-off on LLaVA-OneVision-7B, reducing computational costs to 6.9% of FLOPs while maintaining 99.1% of the original performance. Furthermore, we achieve a 2.28x reduction in Time-To-First-Token (TTFT) and a 1.32x acceleration in decoding throughput, highlighting the practical benefits of our integrated pruning approach for efficient video LLMs inference.

Kele Shao, Keda Tao, Can Qin, Haoxuan You, Yang Sui, Huan Wang• 2025

Related benchmarks

Task	Dataset	Result
Video Understanding	MVBench	Accuracy58.7	635
Video Question Answering	ActivityNet-QA	Accuracy49.9	438
Video Understanding	VideoMME	Score (Overall)64.6	369
Long Video Understanding	LongVideoBench	Score60.4	290
Long Video Understanding	MLVU	--	265
Video Question Answering	VideoMME	Accuracy60.3	254
Video Understanding	MLVU	Score46.7	233
Video Understanding	VideoMME	Overall Score61.6	222
Video Question Answering	MLVU	Accuracy61.6	213
Video Understanding	EgoSchema	EgoSchema Score61.2	185

Showing 10 of 49 rows

Other info

Follow for update

@wizwand_team Discord