Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling
About
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-talker Automatic Speech Recognition | Libri2Mix Clean (test) | WER3.84 | 21 | |
| Target Speaker Extraction | Libri2Mix Clean (test) | -- | 20 | |
| Speaker Verification | Libri2Mix Clean (test) | EER0.39 | 5 | |
| Speech Source Separation | Libri2Mix Clean (test) | SI-SDR (dB)17.3 | 5 | |
| Speech Source Separation | Libri2Mix both (test) | SI-SDR (dB)11.25 | 3 | |
| Target Speaker Extraction | Libri2Mix (test both) | -- | 1 |