Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

About

Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.

Anastasia Zorkina, Alexandr Anikin, Nikita Khmelev, Anastasiya Korenevskaya, Sergey Novoselov, Vladimir Volokhov, Maxim Korenevsky, Yuriy Matveev• 2026

Related benchmarks

TaskDatasetResultRank
Multi-talker Automatic Speech RecognitionLibri2Mix Clean (test)
WER3.84
21
Target Speaker ExtractionLibri2Mix Clean (test)--
20
Speaker VerificationLibri2Mix Clean (test)
EER0.39
5
Speech Source SeparationLibri2Mix Clean (test)
SI-SDR (dB)17.3
5
Speech Source SeparationLibri2Mix both (test)
SI-SDR (dB)11.25
3
Target Speaker ExtractionLibri2Mix (test both)--
1
Showing 6 of 6 rows

Other info

Follow for update