Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

About

Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW's understanding of visual dynamics boosts its downstream performance in visual narrative creation and world simulation, showing improved consistency, controllability, and instruction-following.

Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara Raaghavi Bhattacharya, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji, Li Mi, Syrielle Montariol, Antoine Bosselut• 2026

Related benchmarks

TaskDatasetResultRank
World ModelingLEGO Epic-Kitchen portion (test)
FID10.31
10
Visual Narrative GenerationVinaBench
Non-Character Entity Alignment72.6
9
Visual Narrative GenerationVinaBench (test)
Entity Consistency (Non-Char)77.7
8
World SimulationLEGO Ego4D (test)
FID20.96
4
Image-to-Text SimilarityLEGO Ego4D
BLIP-B Score25.02
3
Image-to-Text SimilarityLEGO Epic-Kitchens
BLIP-B Similarity Score27.27
3
Showing 6 of 6 rows

Other info

Follow for update