XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection

About

Transformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. The code has been publicly released in https://github.com/swagshaw/XLSR-Mamba.

Yang Xiao, Rohan Kumar Das• 2024

Related benchmarks

Task	Dataset	Result
Audio Deepfake Detection	in the wild	EER6.7	65
Audio Deepfake Detection	CodecFake	EER35.26	50
Audio Deepfake Detection	ASVspoof DF 2021	EER1.88	47
Audio Deepfake Detection	ASVspoof LA 2021	EER0.93	41
Audio Deepfake Detection	ASVspoof LA 2019	EER42.1	38
Spoof Speech Detection	ASVspoof LA 2021 (eval)	--	36
Audio Deepfake Detection	FoR	EER6.71	28
Synthetic Speech Detection	ASVspoof DF 2021 (eval)	EER (%)1.88	25
Spoofing Attack Detection	ASVspoof LA 2021	EER0.93	19
Audio Deepfake Detection	ADD 2023 R2	EER20.15	19

Showing 10 of 30 rows

Other info

Code

Follow for update

@wizwand_team Discord