Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

About

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spatiotemporal context, leading to identity drift and poor robustness (e.g., to occlusions), while also inducing lip-shape leakage that degrades lip sync. To bridge this gap, we propose X-Dub, a novel two-stage generative bootstrapping framework leveraging powerful Diffusion Transformers to unlock mask-free dubbing. Our core insight is to repurpose a mask-based inpainting model exclusively as a dedicated data generator to synthesize scalable, high-fidelity pseudo-paired data, which is subsequently utilized to train and bootstrap a robust, mask-free editing model as the final video dubber. The final dubber is liberated from masking artifacts and leverages the complete video input for high-fidelity inference. We further introduce timestep-adaptive multi-phase learning to disentangle conflicting objectives (structure, lip motion, and texture) across diffusion phases, facilitating stable convergence and advanced editing quality. Additionally, we present X-DubBench, a benchmark for diverse scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance with superior lip sync, visual quality, and robustness.

Xu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng, Liyang Chen, Songlin Tang, Jiehui Huang, Xiaoqiang Liu, Pengfei Wan, Zhiyong Wu• 2025

Related benchmarks

TaskDatasetResultRank
Visual DubbingContextDubBench 1.0 (test)
FID9.351
18
Talking Head GenerationTalkVid self-driven 30 clips (held-out)
FVD191.2
16
Lip synchronizationHDTF 52 (test)
Sync-C7.58
12
Visual DubbingUser Study
Realism4.4
9
Visual DubbingHDTF (test)
PSNR34.425
9
Audio-driven Lip SynchronizationHDTF (Cross-identity)
Sync-C Score5.85
8
Lip-sync SynthesisHDTF and TalkVid 30-clip pool (test)
Sync Score4.4
7
Visual DubbingVideo 3-second 25fps 512x512 resolution
Inference Time (s)601
4
Showing 8 of 8 rows

Other info

Follow for update