Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Avocodo: Generative Adversarial Network for Artifact-free Vocoder

About

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency bands, most GAN-based vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we discovered that the multi-scale analysis which focuses on the low-frequency bands causes unintended artifacts, e.g., aliasing and imaging artifacts, which degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based vocoders and propose a GAN-based vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate speech waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band speech waveforms while avoiding aliasing. According to experimental results, Avocodo outperforms baseline GAN-based vocoders, both objectively and subjectively, while reproducing speech with fewer artifacts.

Taejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, Young-Sun Joo• 2022

Related benchmarks

TaskDatasetResultRank
Speech EnhancementSpeech Enhancement (SE) Task (test)
PESQ1.892
22
Neural VocodingLibriTTS (test)
PESQ3.217
18
Speech SynthesisLibriTTS (test)--
17
Speech SynthesisAISHELL3 Mandarin
UTMOS2.126
14
Speech SynthesisSound Effect (evaluation)
M-STFT1.488
13
Neural VocodingLJSpeech 88 (test)
M-STFT1.165
12
Neural VocodingLJSpeech 1.1 (test)
M-STFT1.165
12
Neural VocodingEARS (out-of-domain)
UTMOS2.809
9
Neural VocodingVCTK English Corpus with Unseen Speakers (out-of-domain)
UTMOS3.905
9
VocodingMUSDB18 (test)
ViSQOL (Bass)4.491
8
Showing 10 of 10 rows

Other info

Follow for update