Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents

About

Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail in zero-shot VC scenarios, where models synthesize audio from an audio prompt without training. To address this, we propose VoiceMark, the first zero-shot VC-resistant watermarking method that leverages speaker-specific latents as the watermark carrier, allowing the watermark to transfer through the zero-shot VC process into the synthesized audio. Additionally, we introduce VC-simulated augmentations and VAD-based loss to enhance robustness against distortions. Experiments on multiple zero-shot VC models demonstrate that VoiceMark achieves over 95% accuracy in watermark detection after zero-shot VC synthesis, significantly outperforming existing methods, which only reach around 50%. See our code and demos at: https://huggingface.co/spaces/haiyunli/VoiceMark

Haiyun Li, Zhiyong Wu, Xiaofeng Xie, Jingran Xie, Yaoxun Xu, Hanyang Peng• 2025

Related benchmarks

TaskDatasetResultRank
Audio WatermarkingVCTK, LibriSpeech, and LJSpeech (test)
Detection Accuracy99
96
Watermark RobustnessVCTK, LibriSpeech, and LJSpeech (test)
Accuracy (ACC)99
36
Objective Audio Quality EvaluationLibriSpeech and LJSpeech (test)
NISQA Score4.36
6
Showing 3 of 3 rows

Other info

Follow for update