Dummy Prototypical Networks for Few-Shot Open-Set Keyword Spotting
About
Keyword spotting is the task of detecting a keyword in streaming audio. Conventional keyword spotting targets predefined keywords classification, but there is growing attention in few-shot (query-by-example) keyword spotting, e.g., N-way classification given M-shot support samples. Moreover, in real-world scenarios, there can be utterances from unexpected categories (open-set) which need to be rejected rather than classified as one of the N classes. Combining the two needs, we tackle few-shot open-set keyword spotting with a new benchmark setting, named splitGSC. We propose episode-known dummy prototypes based on metric learning to detect an open-set better and introduce a simple and powerful approach, Dummy Prototypical Networks (D-ProtoNets). Our D-ProtoNets shows clear margins compared to recent few-shot open-set recognition (FSOSR) approaches in the suggested splitGSC. We also verify our method on a standard benchmark, miniImageNet, and D-ProtoNets shows the state-of-the-art open-set detection rate in FSOSR.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Few-shot Audio Classification | FSC-89, NSynth-100, and LS-100 Generalizability Cross-Dataset 5-way 5-shot | Accuracy71.19 | 54 | |
| Few-shot classification | FSC-89 → NSynth-100 | Accuracy70.13 | 31 | |
| Few-shot classification | FSC-89 → LS-100 | Accuracy71.19 | 22 | |
| Few-shot classification | FSC-89 to NSynth-100 | AUROC0.5816 | 18 | |
| Few-shot Open-set Audio Classification | Domestic Environments | Accuracy80.62 | 18 | |
| Few-shot classification | FSC-89 to LS-100 | AUROC63.03 | 9 | |
| Few-shot classification | LS-100 to NSynth-100 | AUROC65.19 | 9 | |
| Few-shot classification | LS-100 → FSC-89 | Accuracy (FSC-89 Few-shot)44.44 | 9 | |
| Few-shot classification | LS-100 → NSynth-100 | Accuracy65.76 | 9 | |
| Open-set Few-shot Classification | NSynth-100 5-way 1-shot | Accuracy86.4 | 9 |