Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

About

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.

Zhiyong Wang, Ruibo Fu, Zhengqi Wen, Jianhua Tao, Xiaopeng Wang, Yuankun Xie, Xin Qi, Shuchen Shi, Yi Lu, Yukun Liu, Chenxing Li, Xuefei Liu, Guanjun Li• 2024

Related benchmarks

TaskDatasetResultRank
Audio Deepfake DetectionASVspoof DF 2021
EER2.54
87
Audio Deepfake Detectionin the wild
EER12.48
76
Audio Deepfake DetectionITW In-the-Wild
EER9.17
51
Audio Deepfake DetectionASVspoof LA 2019
EER74
38
Spoof Speech DetectionASVspoof LA 2021 (eval)--
37
Speech Spoofing DetectionIn-the-Wild (ITW) (eval)
EER9.17
26
Synthetic Speech DetectionASVspoof DF 2021 (eval)
EER (%)2.54
25
Audio Deepfake DetectionASVspoof LA and DF 2021
EER (DF)2.54
22
Deepfake Audio DetectionASVspoof LA 2019
EER (%)74
20
Audio Deepfake DetectionASVspoof LA 2021
EER2.96
12
Showing 10 of 11 rows

Other info

Follow for update