Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration
About
Speech restoration aims to recover clean speech from degraded recordings affected by noise, reverberation, bandwidth reduction, or other distortions, where input and output sampling rates may differ. Existing approaches typically assume matched input-output rates and apply redundant resampling, limiting native multi-rate processing. We formulate this gap as the extended sampling-frequency-independent (xSFI) setting, where a model must operate under decoupled input-output rates, and propose TF-Restormer, a query-based xSFI modeling framework. The model encodes only the observed input band and synthesizes the unobserved high-frequency band through extension queries with band-partitioned cross-attention, yielding an asymmetric encoder-decoder that allocates capacity to analysis while keeping synthesis lightweight. Trained with a perceptual loss, a scaled log-spectral loss, and adversarial supervision via an SFI-STFT discriminator, TF-Restormer attains balanced fidelity-perceptual quality as a single unified model, without redundant resampling across denoising, dereverberation, bandwidth extension, and combined distortion benchmarks under multiple sampling rates.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Speech Enhancement | DNS no-reverb 2020 (test) | -- | 30 | |
| Speech Restoration | DNS Challenge real-recording 2020 | DNSMOS Score3.13 | 14 | |
| Speech Denoising | VCTK+DEMAND | PESQ3.63 | 11 | |
| Speech Restoration | URGENT non-blind 2025 (test) | PESQ2.6 | 11 | |
| Speech Restoration | REVERB real-recording | UTMOS3.14 | 10 | |
| Speech Restoration | DNS with reverb 2020 (test) | SIG Score3.6 | 10 | |
| General Speech Restoration | UNIVERSE 16kHz (test) | UTMOS4.08 | 10 | |
| Speech Restoration | DNS no-reverb 2020 (test) | SIG Score3.65 | 10 | |
| Speech Restoration | URGENT blind 2025 (test) | UTMOS3.37 | 8 | |
| Speech Restoration | REVERB (SimData) | CD2.88 | 8 |