MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

About

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fidelity but suffer from prohibitive computational costs, while GAN-based solutions sacrifice lip-sync accuracy or dental details for real-time performance. We present MuseTalk, a novel two-stage training framework that resolves this trade-off through latent space optimization and spatio-temporal data sampling strategy. Our key innovations include: (1) During the Facial Abstract Pretraining stage, we propose Informative Frame Sampling to temporally align reference-source pose pairs, eliminating redundant feature interference while preserving identity cues. (2) In the Lip-Sync Adversarial Finetuning stage, we employ Dynamic Margin Sampling to spatially select the most suitable lip-movement-promoting regions, balancing audio-visual synchronization and dental clarity. (3) MuseTalk establishes an effective audio-visual feature fusion framework in the latent space, delivering 30 FPS output at 256*256 resolution on an NVIDIA V100 GPU. Extensive experiments demonstrate that MuseTalk outperforms state-of-the-art methods in visual fidelity while achieving comparable lip-sync accuracy. %The codes and models will be made publicly available upon acceptance. The code is made available at \href{https://github.com/TMElyralab/MuseTalk}{https://github.com/TMElyralab/MuseTalk}

Yue Zhang, Zhizhou Zhong, Minhao Liu, Zhaokang Chen, Bin Wu, Yubin Zeng, Chao Zhan, Yingjie He, Junxin Huang, Wenjiang Zhou• 2024

Related benchmarks

Task	Dataset	Result
Talking Head Generation	HDTF (test)	FID27.91	73
Talking Head Generation	HDTF	FID7.25	48
Talking Head Generation	TalkVid self-driven 30 clips (held-out)	FVD136.2	16
Talking Head Generation	CelebV-HQ	FID8.37	15
Lip synchronization	HDTF 52 (test)	Sync-C7.94	12
Talking Face Generation	MEAD Neutral	LSE-C3.55	10
Video-to-Video lip-syncing	TalkVid Self-Reenactment	FID47.78	9
Audio-driven Lip Synchronization	HDTF (Cross-identity)	Sync-C Score6.43	8
Lip synchronization	HDTF	FID8.759	8
Lip synchronization	AIGC-LipSync	FID17.668	8

Showing 10 of 19 rows

Other info

Follow for update

@wizwand_team Discord