HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

About

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated significant improvement in offline video understanding. However, extending these capabilities to streaming video inputs, remains challenging, as existing models struggle to simultaneously maintain stable understanding performance, real-time responses, and low GPU memory overhead. To address this challenge, we propose HERMES, a novel training-free architecture for real-time and accurate understanding of video streams. Based on a mechanistic attention investigation, we conceptualize KV cache as a hierarchical memory framework that encapsulates video information across multiple granularities. During inference, HERMES reuses a compact KV cache, enabling efficient streaming understanding under resource constraints. Notably, HERMES requires no auxiliary computations upon the arrival of user queries, thereby guaranteeing real-time responses for continuous video stream interactions, which achieves 10$\times$ faster TTFT compared to prior SOTA. Even when reducing video tokens by up to 68% compared with uniform sampling, HERMES achieves superior or comparable accuracy across all benchmarks, with up to 11.4% gains on streaming datasets.

Haowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng, Xipeng Qiu• 2026

Related benchmarks

Task	Dataset	Result
Video Understanding	MVBench	Accuracy65.53	635
Visual Question Answering	GQA	Accuracy57.6	524
Optical Character Recognition	OCRBench	--	486
Visual Question Answering	RealworldQA	Accuracy50.2	327
Streaming Video Understanding	StreamingBench	Overall79.44	308
Long Video Understanding	LVBench	Accuracy49.8	267
Text-based Visual Question Answering	TextVQA	TextVQA Accuracy58	141
Real-Time Visual Understanding	StreamingBench	Overall Score79.44	134
Long Video Understanding	Video-MME Long	Accuracy65	120
Visual Question Answering	MMVP	Accuracy45.3	82

Showing 10 of 28 rows

Other info

GitHub

Follow for update

@wizwand_team Discord