End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection

About

Artefacts that serve to distinguish bona fide speech from spoofed or deepfake speech are known to reside in specific subbands and temporal segments. Various approaches can be used to capture and model such artefacts, however, none works well across a spectrum of diverse spoofing attacks. Reliable detection then often depends upon the fusion of multiple detection systems, each tuned to detect different forms of attack. In this paper we show that better performance can be achieved when the fusion is performed within the model itself and when the representation is learned automatically from raw waveform inputs. The principal contribution is a spectro-temporal graph attention network (GAT) which learns the relationship between cues spanning different sub-bands and temporal intervals. Using a model-level graph fusion of spectral (S) and temporal (T) sub-graphs and a graph pooling strategy to improve discrimination, the proposed RawGAT-ST model achieves an equal error rate of 1.06 % for the ASVspoof 2019 logical access database. This is one of the best results reported to date and is reproducible using an open source implementation.

Hemlata Tak, Jee-weon Jung, Jose Patino, Madhu Kamble, Massimiliano Todisco, Nicholas Evans• 2021

Related benchmarks

Task	Dataset	Result
Audio Deepfake Detection	in the wild	EER37.8	76
Audio Deepfake Detection	ITW In-the-Wild	EER33.8	51
Audio Deepfake Detection	ASVspoof 2021	EER22.3	39
Audio Deepfake Detection	ASVspoof LA 2019	--	38
Audio Deepfake Detection	ASVspoof 2019	EER3	37
Audio Spoofing Detection	ASVspoof Logical Access 2019 (Evaluation)	EER1.06	30
Audio Deepfake Detection	FoR	--	28
Speech Deepfake Detection	ASVspoof logical access (LA) 2019 (eval)	min-tDCF0.0335	21
Audio Deepfake Detection	MLAAD-EN	EER33.9	18
Audio Deepfake Detection	WaveFake	Accuracy68.4	15

Showing 10 of 35 rows

Other info

Code

Follow for update

@wizwand_team Discord