Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

About

Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of video tokens and key-value (KV) caches. We present OmniMem, a memory-efficient streaming framework designed specifically for audio-visual LLMs. Unlike existing compression methods that treat all tokens uniformly, OmniMem introduces a modality-aware memory allocation strategy that separately manages visual and audio contexts, addressing the severe token imbalance between the two modalities. OmniMem further preserves informative and non-redundant KV states through perturbation-aware memory selection, enabling compact memory without sacrificing long-range understanding. To strengthen compression under realistic deployment constraints, we also explore budget-aware fine-tuning, which encourages the model to consolidate useful information into retained memory. Experiments on VideoMME Long, LVBench, and LVOmniBench with video-SALMONN 2+ and Qwen-2.5-Omni show that OmniMem consistently improves over strong training-free compression baselines by 2-4% absolute accuracy under the same memory budgets, with an additional 1-2% gain after fine-tuning.

Guangzhi Sun, Yixuan Li, Yudong Yang, Chao Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Long Video UnderstandingLVBench
Accuracy55.7
267
Long Video UnderstandingVideo-MME Long
Accuracy70.2
120
Long Video UnderstandingLV-Omni-Bench
Accuracy43.1
17
Streaming Video UnderstandingStreamingBench Realtime partition
Accuracy78.5
9
Streaming Video UnderstandingStreamingBench Omni-source partition
Accuracy60.9
9
Streaming Video UnderstandingStreamingBench Contextual partition
Accuracy40.7
9
Showing 6 of 6 rows

Other info

Follow for update