Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

About

To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled mechanism to control the utility-privacy trade-off, operating at fixed design points. We propose DiffAnon, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation. DiffAnon refines acoustic detail over semantic embeddings of an RVQ codec, enabling smooth interpolation between anonymization strength and prosodic fidelity within a single model. To the best of our knowledge, it is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control. Experiments demonstrate structured trade-off behavior, achieving strong utility while maintaining competitive privacy across controllable operating points.

Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn, Berrak Sisman• 2026

Related benchmarks

TaskDatasetResultRank
Voice AnonymizationIEMOCAP (dev)
UAR52.57
18
Voice AnonymizationIEMOCAP (test)
UAR50.8
18
Voice AnonymizationLibriSpeech (dev)
WER4.91
18
Voice AnonymizationLibriSpeech (test)
WER4.62
18
Showing 4 of 4 rows

Other info

Follow for update