Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor

About

We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based attractor (TDA) calculation module that can deal with an unknown number of speakers, and 3) triple-path processing blocks that can model inter-speaker relations. Given a fixed, small set of learned speaker queries and the mixture embedding produced by the dual-path blocks, TDA infers the relations of these queries and generates an attractor vector for each speaker. The estimated attractors are then combined with the mixture embedding by feature-wise linear modulation conditioning, creating a speaker dimension. The mixture embedding, conditioned with speaker information produced by TDA, is fed to the final triple-path blocks, which augment the dual-path blocks with an additional pathway dedicated to inter-speaker processing. The proposed approach outperforms the previous best reported in the literature, achieving 24.0 and 23.7 dB SI-SDR improvement (SI-SDRi) on WSJ0-2 and 3mix respectively, with a single model trained to separate 2- and 3-speaker mixtures. The proposed model also exhibits strong performance and generalizability at counting sources and separating mixtures with up to 5 speakers.

Younglo Lee, Shukjae Choi, Byeong-Yeol Kim, Zhong-Qiu Wang, Shinji Watanabe• 2024

Related benchmarks

Task	Dataset	Result
Speech Separation	WSJ0-2Mix (test)	SDRi (dB)23.9	160
Speech Separation	WSJ0 3mix	SI-SNRi23.7	17
Monaural Speech Separation	WSJ0-2Mix	ΔSI-SDR (dB)24	13
Monaural Speech Separation	WSJ0 3mix	ΔSI-SDR (dB)23.7	13
Speech Separation	WSJ0 4mix	SI-SNRi22	9
Speech Separation	WSJ0 5mix	SI-SNRi21	8
Monaural Speech Separation	WSJ0 4mix	Delta SI-SDR (dB)22	7
Monaural Speech Separation	WSJ0 5mix	ΔSI-SDR (dB)21	6

Showing 8 of 8 rows

Other info

Code

Follow for update

@wizwand_team Discord