Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

About

Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. However, advanced video agents often rely on ``detective-style'' iterative reasoning for action control (e.g., $\mathtt{search}$) and evidence aggregation, incurring prohibitive costs and latency. We argue that such heavy reasoning primarily compensates for the lack of global context and semantic misalignment in retrieval. This paper introduces Light-Omni, a multimodal agent framework for reflexive and lightweight video understanding. It achieves this through dual contextual states that instantly build the required context in a single forward pass. First, we maintain a global state, a finite-sized multimodal script continuously consolidated from episodic memory, serving as the global context for Light-Omni. Through hierarchical merging, it preserves recent details while summarizing past events. Second, conditioned on this global context, we generate a parametric latent state that directly drives autonomous actions and produces retrieval embeddings, with minimal latency. Benefiting from this coupled design, Light-Omni achieves semantically aligned retrieval and reflexive responses while avoiding iterative reasoning. Extensive experiments validate the effectiveness of Light-Omni across multiple video benchmarks. Notably, it outperforms M3-Agent with an average 2.4% accuracy gain, a 12.1$\times$ speedup, and a 2.6$\times$ improvement in GPU memory efficiency. Furthermore, it serves as a memory system to enhance both the performance and efficiency of existing MLLMs. Project page: https://clare-nie.github.io/Light-Omni.

Chang Nie, Jiaju Wei, Junlan Feng, Chaoyou Fu, Caifeng Shan• 2026

Related benchmarks

TaskDatasetResultRank
Long Video UnderstandingLVBench
Accuracy49.9
267
Long Video UnderstandingVideo-MME Long
Accuracy66.1
120
Long Video UnderstandingVideoMME Long
Accuracy69
15
Long Video UnderstandingLVBench
Accuracy50.1
15
Long Video Question AnsweringVideoMME Long
Accuracy (Long)66.1
14
Long Video UnderstandingHippoVlog
Accuracy78.5
11
Long Video UnderstandingVideo-MME LVBench HippoVlog long
Average Accuracy64.8
11
Showing 7 of 7 rows

Other info

GitHub

Follow for update