HLTCOE JHU Submission to the Voice Privacy Challenge 2024
About
We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, Leibny Paola Garc\'ia-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner• 2024
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Voice Anonymization | LibriSpeech (dev) | WER3.45 | 18 | |
| Voice Anonymization | LibriSpeech (test) | WER3.19 | 18 | |
| Voice Anonymization | IEMOCAP (test) | UAR47.1 | 18 | |
| Voice Anonymization | IEMOCAP (dev) | UAR47.07 | 18 |
Showing 4 of 4 rows