Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Speaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

About

Sophisticated generative speech technology can undermined the reliability of voice biometrics. While spoofing detection systems excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. In this paper, we show that such issues could be caused by speaker bias, where models learn individual voice traits rather than markers of manipulation or generation. We propose a teacher-student framework for speaker-invariant spoofing detection that disentangles identity without requiring speaker labels. We leverage a pre-trained speaker recognition teacher to guide a student model via a gradient reversal layer. To control the balance between suppressing cues related to voice identity with the preservation of those related to spoofing detection, we integrate a Variational Information Bottleneck. Evaluations across nine datasets show our model achieves a 25.7% relative reduction to the EER compared to the MHFA baseline.

Anh-Tuan Dao, Driss Matrouf, Mickael Rouvier, Nicholas Evans• 2026

Related benchmarks

TaskDatasetResultRank
Spoofing Attack DetectionASVspoof LA 2021
EER6.12
37
Spoofing Attack DetectionASVspoof DF 2021
EER3.71
31
Speech Spoofing DetectionIn-the-Wild (ITW) (eval)
EER2.31
26
Anti-spoofingPooled
EER10.15
23
Anti-spoofingITW
EER2.31
21
Spoofing DetectionASVspoof 5 (eval)
EER5.19
18
Spoofing DetectionASVspoof 2019 (eval)
EER5.18
13
Audio Spoof DetectionASVspoof LA 2019
A075.18
11
Spoofing DetectionASVspoof DeepFake 2021 (test)
EER3.71
7
Spoofing DetectionFoR Fake-or-Real (evaluation)
EER5.16
7
Showing 10 of 14 rows

Other info

Follow for update