From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
About
Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spatiotemporal context, leading to identity drift and poor robustness (e.g., to occlusions), while also inducing lip-shape leakage that degrades lip sync. To bridge this gap, we propose X-Dub, a novel two-stage generative bootstrapping framework leveraging powerful Diffusion Transformers to unlock mask-free dubbing. Our core insight is to repurpose a mask-based inpainting model exclusively as a dedicated data generator to synthesize scalable, high-fidelity pseudo-paired data, which is subsequently utilized to train and bootstrap a robust, mask-free editing model as the final video dubber. The final dubber is liberated from masking artifacts and leverages the complete video input for high-fidelity inference. We further introduce timestep-adaptive multi-phase learning to disentangle conflicting objectives (structure, lip motion, and texture) across diffusion phases, facilitating stable convergence and advanced editing quality. Additionally, we present X-DubBench, a benchmark for diverse scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance with superior lip sync, visual quality, and robustness.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Visual Dubbing | ContextDubBench 1.0 (test) | FID9.351 | 18 | |
| Talking Head Generation | TalkVid self-driven 30 clips (held-out) | FVD191.2 | 16 | |
| Lip synchronization | HDTF 52 (test) | Sync-C7.58 | 12 | |
| Visual Dubbing | User Study | Realism4.4 | 9 | |
| Visual Dubbing | HDTF (test) | PSNR34.425 | 9 | |
| Audio-driven Lip Synchronization | HDTF (Cross-identity) | Sync-C Score5.85 | 8 | |
| Lip-sync Synthesis | HDTF and TalkVid 30-clip pool (test) | Sync Score4.4 | 7 | |
| Visual Dubbing | Video 3-second 25fps 512x512 resolution | Inference Time (s)601 | 4 |