Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

About

Most neural vocoders employ band-limited mel-spectrograms to generate waveforms. If full-band spectral features are used as the input, the vocoder can be provided with as much acoustic information as possible. However, in some models employing full-band mel-spectrograms, an over-smoothing problem occurs as part of which non-sharp spectrograms are generated. To address this problem, we propose UnivNet, a neural vocoder that synthesizes high-fidelity waveforms in real time. Inspired by works in the field of voice activity detection, we added a multi-resolution spectrogram discriminator that employs multiple linear spectrogram magnitudes computed using various parameter sets. Using full-band mel-spectrograms as input, we expect to generate high-resolution signals by adding a discriminator that employs spectrograms of multiple resolutions as the input. In an evaluation on a dataset containing information on hundreds of speakers, UnivNet obtained the best objective and subjective results among competing models for both seen and unseen speakers. These results, including the best subjective score for text-to-speech, demonstrate the potential for fast adaptation to new speakers without a need for training from scratch.

Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, Juntae Kim• 2021

Related benchmarks

TaskDatasetResultRank
Music Source SeparationMUSDB18 HQ (test)
SDR (Drums)4.23
48
Waveform GenerationMUSDB18 out-of-distribution vocal samples HQ (test)
M-STFT1.1377
19
Audio GenerationLibriTTS (dev)
M-STFT0.8959
18
Speech SynthesisLibriTTS (test)
MOS4.8042
17
Waveform GenerationLibriTTS 24,000 Hz (test)
UTMOS3.5233
13
Waveform GenerationLibriTTS (dev)
M-STFT0.8947
12
Speech SynthesisLibriTTS 24,000 Hz (test)
MOS4.2
11
Audio SynthesisLJSpeech (unseen)
MAE0.3356
10
Neural VocodingLibriTTS clean (dev)
MAE0.2772
10
Neural VocodingVCTK 100 audio clips (unseen)
MAE0.2753
10
Showing 10 of 16 rows

Other info

Follow for update