Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

About

Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensive models, which limit their practical use. The study introduces the ExpSpeech-Net deepfake detection (SqN-R-DFD) model, which utilizes SqueezeNet and RNN (Recurrent Neural Network) as its backbone, providing a lightweight and efficient deepfake detection framework that simultaneously analyzes facial expressions and speech patterns. The approach incorporates advanced feature extraction, such as ISLBT-based features for image and MPNCC for signals, along with a smart feature-selection strategy using SASMA (Sandpiper-Assisted Slime Mould Algorithm), ensuring optimal and balanced input to the detection models. By combining SqueezeNet and an RNN, subtle inconsistencies in deepfake videos are captured effectively. The framework achieves 94.5% accuracy, precision of 99.3%, and F-measure of 96.8%, outperforming conventional methods. This demonstrates that integrating multiple modalities with intelligent preprocessing and feature selection enables practical, real-time deepfake detection suitable for everyday applications.

Ruchika Sharma, Rudresh Dwivedi• 2026

Related benchmarks

TaskDatasetResultRank
Deepfake DetectionDeepfakeTIMIT Dataset2 (K-Fold Cross Validation)
Accuracy97.7
40
Deepfake DetectionWLDR Dataset1 (Fold 2)
Accuracy92.7
8
Deepfake DetectionWLDR Fold 3 Dataset1 (K-Fold Cross Validation)
Accuracy94.9
8
Deepfake DetectionWLDR Fold 4 Dataset1 (K-Fold Cross Validation)
Accuracy95.4
8
Deepfake DetectionWLDR K-Fold Cross Validation Dataset1 (Fold 5)
Accuracy95.5
8
Deepfake DetectionWLDR Dataset1 (Fold 6)
Accuracy96.5
8
Showing 6 of 6 rows

Other info

Follow for update