Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

About

While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing methods suffer from temporal misalignment, limited controllability, and incomplete scripting. In this paper, we propose MAVIN, the first framework for multi-shot audio-visual generation with customized narrative control. To resolve temporal misalignment, we propose boundary-aware attention, which leverages hierarchical captions and boundary-aware token routing to render audio-visual elements within their respective temporal boundaries. To improve the controllability for multi-subject scenarios, we propose ID-aware propagation, utilizing identity embeddings and an identity-aware mask to bind specific identities to consistent visual appearances and vocal timbres. To provide comprehensive audio-visual narratives, we present a multi-agent scripting pipeline to transform free-form user inputs into hierarchical captions. Furthermore, we construct MAVINSet, a multi-shot audio-visual dataset for robust training and evaluation. Extensive experiments demonstrate that MAVIN achieves state-of-the-art performance, opening up a new avenue for integrating generative models into professional filmmaking workflows.

Kaiqi Liu, Yunyao Mao, Ziqi Cai, Zheng Geng, Jing Wang, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Shuchen Weng, Boxin Shi• 2026

Related benchmarks

TaskDatasetResultRank
Multi-shot Audio-Visual GenerationMAVINSet subjective 20 samples
AVQ36.8
11
Multi-shot Audio-Visual GenerationMAVINSet high-fidelity benchmark 1K-sample (test)
FVD231.6
11
Personalized Audiovisual GenerationIdentity-Referenced Audiovisual Generation Dataset S2 (test)
FVD241.9
5
Text-to-Audio-Video GenerationMAVINSet 1.0 (test)
FVD231.6
3
Showing 4 of 4 rows

Other info

Follow for update