Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

About

Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearing technologies. However, our analysis reveals that high within-trial performance can be driven by trial-specific EEG structure that acts as shortcuts for target selection, leading to poor generalization on unseen trials. To overcome this gap, we propose TRUST-TSE, a two-stage framework to mitigate shortcut learning. By introducing contrastive pretraining with attended-speaker negative sampling, we encourage the EEG encoder to capture fine-grained EEG--speech alignment while suppressing trial-identity cues. We also employ a confidence-weighted extraction objective based on EEG--source similarity to guide extraction using the learned representations. Experiments on KUL and DTU datasets show that TRUST-TSE outperforms end-to-end baselines under strict cross-trial protocols, addressing a key reliability bottleneck of existing approaches.

Wonchul Shin, Inyong Choi, Kyogu Lee• 2026

Related benchmarks

TaskDatasetResultRank
Target SelectionKUL (cross-trial)
Selection Accuracy65.62
5
Target SelectionDTU (cross-trial)
Selection Accuracy71.58
5
EEG-guided Target Speaker ExtractionKUL 64-channel (test)
Accuracy62.27
3
EEG-guided Target Speaker ExtractionDTU 64-channel (test)
Accuracy70.4
3
Target Speech ExtractionKUL (unseen subjects)
Accuracy64.31
3
Target Speech ExtractionDTU (unseen subjects)
Accuracy67.16
3
Showing 6 of 6 rows

Other info

Follow for update