Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck
About
Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adversarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Spoofing Attack Detection | ASVspoof LA 2021 | EER5.58 | 37 | |
| Spoofing Attack Detection | ASVspoof DF 2021 | EER3.09 | 31 | |
| Anti-spoofing | Pooled | EER8.72 | 23 | |
| Anti-spoofing | ITW | EER1.88 | 21 | |
| Audio anti-spoofing | ASVspoof 5 (evaluation) | EER5.26 | 17 | |
| Spoofing Detection | ASVspoof 2019 (eval) | EER4.07 | 13 | |
| Audio Spoof Detection | ASVspoof LA 2019 | -- | 11 | |
| Audio anti-spoofing | in the wild | EER1.88 | 7 | |
| Spoofing Detection | FoR | EER3.35 | 6 | |
| Spoofing Detection | CodecFake | EER20.28 | 6 |