Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Latent Visual Cache for Video Reasoning

About

Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visual grounding weakens during generation. This phenomenon has been widely observed and is known as Visual Anchoring Decay. To fill this gap, we introduce Latent Video Cache (Latent-VC), a recurrent latent visual cache inserted into the decoder to preserve compact visual memories throughout reasoning. The cache is trained with supervised contrastive cache alignment and vision-grounded GRPO with a latent grounding reward, while maintaining strict train-inference alignment through native decoder hidden states. Built on Qwen3.5-9B, Latent-VC consistently outperforms strong CoT and SFT+GRPO baselines across six video benchmarks, with especially clear gains on grounding-intensive and long-video tasks. In addition, it also achieves higher accuracy with substantially shorter responses, suggesting that latent visual caching improves video reasoning by preserving visual evidence rather than relying on longer textual chains.

Yongheng Zhang, Zhipeng Xu, Hao Wu, Yinghui Li, Di Yin, Xing Sun, Philip S. Yu• 2026

Related benchmarks

TaskDatasetResultRank
Video ReasoningVideoMMMU
Accuracy65.4
141
Video ReasoningVSI-Bench
Accuracy61.9
101
Video ReasoningVideo-MME
Accuracy66.1
73
Video ReasoningMMVU
Accuracy68.6
71
Video ReasoningMVBench--
56
Video ReasoningTempCompass
Score72.6
46
Showing 6 of 6 rows

Other info

Follow for update