Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

HLTCOE JHU Submission to the Voice Privacy Challenge 2024

About

We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.

Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, Leibny Paola Garc\'ia-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner• 2024

Related benchmarks

TaskDatasetResultRank
Voice AnonymizationLibriSpeech (dev)
WER3.45
18
Voice AnonymizationLibriSpeech (test)
WER3.19
18
Voice AnonymizationIEMOCAP (test)
UAR47.1
18
Voice AnonymizationIEMOCAP (dev)
UAR47.07
18
Showing 4 of 4 rows

Other info

Follow for update