Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VideoMemory: Toward Consistent Video Generation via Memory Integration

About

Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality short clips but often fail to preserve entity identity and appearance when scenes change or when entities reappear after long temporal gaps. We present VideoMemory, an entity-centric framework that integrates narrative planning with visual generation through a Dynamic Memory Bank. Given a structured script, a multi-agent system decomposes the narrative into shots, retrieves entity representations from memory, and synthesizes keyframes and videos conditioned on these retrieved states. The Dynamic Memory Bank stores explicit visual and semantic descriptors for characters, props, and backgrounds, and is updated after each shot to reflect story-driven changes while preserving identity. This retrieval-update mechanism enables consistent portrayal of entities across distant shots and supports coherent long-form generation. To evaluate this setting, we construct a 54-case multi-shot consistency benchmark covering character-, prop-, and background-persistent scenarios. Extensive experiments show that VideoMemory achieves strong entity-level coherence and high perceptual quality across diverse narrative sequences.

Jinsong Zhou, Yihua Du, Xinli Xu, Luozhou Wang, Zijie Zhuang, Yehang Zhang, Shuaibo Li, Xiaojun Hu, Bolan Su, Ying-cong Chen• 2026

Related benchmarks

TaskDatasetResultRank
Video GenerationVBench
Subject Consistency92.07
8
Video GenerationVBench-Long 8 continuous scenes for 1-min videos
Semantic Alignment0.1926
8
Video SynthesisLVbench-C 3-min, 24 scenes
Semantic Alignment20.62
8
Video SynthesisLVbench-C 5-min, 40 scenes
Semantic Alignment Score15.45
8
Long Video SynthesisVBench Long
Character Consistency4.4
7
Multi-shot Video Consistency54 Multi-shot Benchmark Cases 1.0 (full)
Character Consistency (4 shots)0.61
6
Multi-shot Video Consistency54 multi-shot benchmark cases 1.0 (test)
Character Consistency95.83
5
Showing 7 of 7 rows

Other info

Follow for update