Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DuraMark: Duration-Embedded Watermarking in LLM-based TTS

About

Large language model (LLM)-based text-to-speech (TTS) models have achieved remarkable voice cloning capabilities, raising concerns about potential deepfake misuse. Speech watermarking mitigates this by embedding traceable information into generated speech. Mainstream watermarking methods operate at the signal level (waveform or spectrogram), rendering the watermark vulnerable to generative attacks (e.g., neural codec and vocoder). To address this, we propose DuraMark, a robust information-level watermarking framework. It utilizes syllable duration editing to achieve watermark embedding. Specifically, DuraMark integrates a duration-controllable LLM-based TTS model to edit syllable durations during synthesis, coupled with a duration extractor to extract these durations for detection. Experiments demonstrate DuraMark's superior robustness against generative attacks, significantly outperforming signal-level baselines. Audio samples are available at https://muzw.github.io/duramark_demo/.

Zhenwei Mou, Weili Jiang, Liping Chen, Zhen-Hua Ling, Kong Aik Lee, Kai Gao, Boyu Zhao• 2026

Related benchmarks

TaskDatasetResultRank
Speech Watermarking RobustnessAISHELL-3 (test)
TPR99.9
100
Speech Quality EvaluationAISHELL-3 (test)
CER8.54
6
Showing 2 of 2 rows

Other info

Follow for update