Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Improved DeepFake Detection Using Whisper Features

About

With a recent influx of voice generation methods, the threat introduced by audio DeepFake (DF) is ever-increasing. Several different detection methods have been presented as a countermeasure. Many methods are based on so-called front-ends, which, by transforming the raw audio, emphasize features crucial for assessing the genuineness of the audio sample. Our contribution contains investigating the influence of the state-of-the-art Whisper automatic speech recognition model as a DF detection front-end. We compare various combinations of Whisper and well-established front-ends by training 3 detection models (LCNN, SpecRNet, and MesoNet) on a widely used ASVspoof 2021 DF dataset and later evaluating them on the DF In-The-Wild dataset. We show that using Whisper-based features improves the detection for each model and outperforms recent results on the In-The-Wild dataset by reducing Equal Error Rate by 21%.

Piotr Kawa, Marcin Plata, Micha{\l} Czuba, Piotr Szyma\'nski, Piotr Syga• 2023

Related benchmarks

TaskDatasetResultRank
Audio Deepfake DetectionCodecFake
EER7.92
50
Speech Deepfake DetectionCodecFake (CF) seen setting
Accuracy94.41
32
Speech Deepfake DetectionSeaCF (seen setting)
Accuracy87.69
32
Speech Deepfake DetectionICF
Accuracy91.98
23
Showing 4 of 4 rows

Other info

Follow for update