Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

About

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual cues, to reconstruct incomplete or masked speech signals in noisy environments. However, existing speech enhancement (SE) approaches often overlook the rich semantic information embedded in speech, which is crucial for improving intelligibility, speaker consistency, and overall quality of enhanced speech signals. To enrich the SE model with semantic information, we employ language models as an efficient semantic learner and propose a comprehensive framework tailored for language model-based speech enhancement, called \textit{GenSE}. Specifically, we approach SE as a conditional language modeling task rather than a continuous signal regression problem defined in existing works. This is achieved by tokenizing speech signals into semantic tokens using a pre-trained self-supervised model and into acoustic tokens using a custom-designed single-quantizer neural codec model. To improve the stability of language model predictions, we propose a hierarchical modeling method that decouples the generation of clean semantic tokens and clean acoustic tokens into two distinct stages. Moreover, we introduce a token chain prompting mechanism during the acoustic token generation stage to ensure timbre consistency throughout the speech enhancement process. Experimental results on benchmark datasets demonstrate that our proposed approach outperforms state-of-the-art SE systems in terms of speech quality and generalization capability.

Jixun Yao, Hexin Liu, Chen Chen, Yuchen Hu, EngSiong Chng, Lei Xie• 2025

Related benchmarks

TaskDatasetResultRank
Speech EnhancementDNS no-reverb 2020 (test)
Signal Score (SIG)3.65
30
Speech EnhancementDNS Challenge Without Reverb (test)
SIG Score3.65
26
Personalized Speech EnhancementDNS Track 1: Headset 5 (test)
SIG Score4.13
19
Personalized Speech EnhancementDNS Track 2: Speakerphone Blind 5 (test)
SIG Score3.92
19
Speech EnhancementDNS blind synthetic with reverb 2020 (test)
SIG Score3.51
16
Speech EnhancementDNS blind (real recordings) 2020 (test)
SIG Score3.1
16
Speech RestorationDNS Challenge With Reverb 2020 (test)
SIG Score3.49
14
Noise SuppressionInterspeech DNS Challenge blind No Reverb 2020 (test)
SIG Score3.65
10
Speech RestorationDNS no-reverb 2020 (test)
SIG Score3.65
10
Noise SuppressionInterspeech DNS Challenge With Reverb 2020 (test)
SIG Score3.49
10
Showing 10 of 14 rows

Other info

Follow for update