Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

KFC-KWS: Keyframe Fusion with CTC for User-Defined Keyword Spotting

About

User-defined keyword spotting (KWS) enables personalized voice interaction by detecting user-specified keywords. A key challenge in this task is distinguishing target keywords from phonetically confusable alternatives. To address this challenge, we propose KFC-KWS, a multimodal framework that leverages connectionist temporal classification (CTC)-guided keyframe selection. Specifically, we exploit the peaky posterior distributions of CTC to identify high-confidence phoneme frames, enabling precise alignment across audio, phoneme, and text modalities. These keyframes are then fused with full-utterance representations through cross-attention to capture both local discriminative cues and global contextual information. On LibriPhrase, KFC-KWS achieves the best-balanced performance (98.73% AUC) and substantially outperforms advanced baselines on the challenging hard subset (97.65% AUC and 7.75% EER), demonstrating its effectiveness in discriminating between highly confusable keywords.

Jin Li, Wenbin Jiang, Ji Hu• 2026

Related benchmarks

TaskDatasetResultRank
Keyword SpottingLibriPhrase Easy (LPE)
EER1.94
51
Keyword SpottingLibriPhrase Hard (LPH)
EER0.0775
25
User-defined keyword spottingLibriPhrase LPH hard
AUC96.54
9
User-defined keyword spottingLibriPhrase Balanced (combined)
AUC98.06
9
User-defined keyword spottingLibriPhrase LPE (easy)
AUC99.58
9
Keyword SpottingLibriPhrase (LPH + LPE)/2 (balanced)
AUC98.73
5
Showing 6 of 6 rows

Other info

Follow for update