Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection

About

This paper describes our submitted systems to the ASVspoof 5 Challenge Track 1: Speech Deepfake Detection - Open Condition, which consists of a stand-alone speech deepfake (bonafide vs spoof) detection task. Recently, large-scale self-supervised models become a standard in Automatic Speech Recognition (ASR) and other speech processing tasks. Thus, we leverage a pre-trained WavLM as a front-end model and pool its representations with different back-end techniques. The complete framework is fine-tuned using only the trained dataset of the challenge, similar to the close condition. Besides, we adopt data-augmentation by adding noise and reverberation using MUSAN noise and RIR datasets. We also experiment with codec augmentations to increase the performance of our method. Ultimately, we use the Bosaris toolkit for score calibration and system fusion to get better Cllr scores. Our fused system achieves 0.0937 minDCF, 3.42% EER, 0.1927 Cllr, and 0.1375 actDCF.

Theophile Stourbe, Victor Miara, Theo Lepage, Reda Dehak• 2024

Related benchmarks

TaskDatasetResultRank
Spoofing Attack DetectionASVspoof LA 2021
EER6.8
37
Spoofing Attack DetectionASVspoof DF 2021
EER4.44
31
Anti-spoofingPooled
EER5.48
23
Anti-spoofingITW
EER2.27
21
Audio anti-spoofingASVspoof DF 2021 (hidden)
EER3.53
19
Audio anti-spoofingASVspoof LA 2021 (hidden)
EER5.86
19
Audio anti-spoofingASVspoof 5 (evaluation)
EER5.56
17
Fake DetectionASVspoof5 (eval)
EER4.64
16
Spoofing DetectionASVspoof 2019 (eval)
EER9.48
13
Audio anti-spoofingWild
EER0.0145
12
Showing 10 of 18 rows

Other info

Follow for update