Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

About

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the complex-valued convolutional layers. In this paper, we propose a complex convolutional block attention module (CCBAM) to boost the representation power of the complex-valued convolutional layers by constructing more informative features. The CCBAM is a lightweight and general module which can be easily integrated into any complex-valued convolutional layers. We integrate CCBAM with the deep complex U-Net and CRN to enhance their performance for speech enhancement. We further propose a mixed loss function to jointly optimize the complex models in both time-frequency (TF) domain and time domain. By integrating CCBAM and the mixed loss, we form a new end-to-end (E2E) complex speech enhancement framework. Ablation experiments and objective evaluations show the superior performance of the proposed approaches (https://github.com/modelscope/ClearerVoice-Studio).

Shengkui Zhao, Trung Hieu Nguyen, Bin Ma• 2021

Related benchmarks

TaskDatasetResultRank
Speech EnhancementDNS no-reverb 2020 (test)--
30
Speech RestorationDNS no-reverb 2020 (test)
SIG Score3.58
10
Speech RestorationDNS with reverb 2020 (test)
SIG Score2.93
10
Speech EnhancementDNS dataset
PESQ (Wideband)3.23
9
Speech EnhancementDataset-2 DNS Challenge Official synthetic 2020 (test)
PESQ3.3
8
Speech EnhancementDNS Challenge official synthetic dataset-2 (test)
PESQ3.3
8
Speech EnhancementDataset-1 WSJ0, DEMAND, RNNoise 1.0 (test)
PESQ3.44
7
Showing 7 of 7 rows

Other info

Follow for update