Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Bridging Video Understanding and Generation in a Unified Framework

About

Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video understanding favors compact, discriminative semantic representations, whereas video generation requires dense signals that preserve visual details and temporal coherence. Videos naturally capture both spatial semantics and temporal dynamics, making them a more suitable modality for unified multimodal modeling compared to static images. In this paper, we propose Vega, a unified framework that bridges video understanding and generation. Vega leverages a shared vocabulary to jointly model text and visual representations and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering. Specifically, the AR model focuses on predicting semantically meaningful visual tokens for keyframes, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames. Extensive experiments demonstrate that Vega achieves strong performance on video generation benchmarks such as VBench and video understanding benchmarks like VideoMME.

Yuqi Wang, Runyi Li, Ruoyu Feng, Renjie Chen, Wenfeng Lin, Mingyu Guo• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Video GenerationVBench
Quality Score82.64
209
Video UnderstandingVideoMME
Score (%)69.4
28
Video UnderstandingLongVideoBench (full)
LongVideoBench Score60.1
10
Video UnderstandingEgoSchema fullset
EgoSchema Accuracy52.9
9
Image-to-Video GenerationVBench++
Total Score86.7
8
Video UnderstandingMLVU (full)
MLVU Score71.6
7
Video UnderstandingNextQA (full)
NextQA Accuracy76.3
6
Showing 7 of 7 rows

Other info

Follow for update