Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Post-Training Speech Enhancement Language Models with Perceptual Rewards

About

Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entropy rather than the perceptual metrics used for evaluation. We introduce a post-training stage for autoregressive speech enhancement language models using Group Sequence Policy Optimization (GSPO) with multi-metric perceptual rewards. Our method directly optimizes non-differentiable quality metrics (DNSMOS, WER, and UTMOS) as reward signals, without learned surrogates or offline preference pairs. Applied to two autoregressive base models, UniSE and GenSE, our approach achieves state-of-the-art results on the DNS2020 benchmark. A human evaluation ablation further shows that the composite multi-metric reward is preferred over any single-metric variant, confirming that multi-reward optimization avoids the reward hacking observed with single-metric training.

Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Antonis Asonitis, Roger Wattenhofer• 2026

Related benchmarks

TaskDatasetResultRank
Speech EnhancementDNS no-reverb 2020 (test)
Signal Score (SIG)3.75
30
Personalized Speech EnhancementDNS Track 1: Headset 5 (test)
SIG Score4.75
19
Personalized Speech EnhancementDNS Track 2: Speakerphone Blind 5 (test)
SIG Score4.73
19
Speech EnhancementDNS blind synthetic with reverb 2020 (test)
SIG Score3.76
16
Speech EnhancementDNS blind (real recordings) 2020 (test)
SIG Score3.63
16
Showing 5 of 5 rows

Other info

Follow for update