Few-shot Class-incremental Audio Classification Using Adaptively-refined Prototypes
About
New classes of sounds constantly emerge with a few samples, making it challenging for models to adapt to dynamic acoustic environments. This challenge motivates us to address the new problem of few-shot class-incremental audio classification. This study aims to enable a model to continuously recognize new classes of sounds with a few training samples of new classes while remembering the learned ones. To this end, we propose a method to generate discriminative prototypes and use them to expand the model's classifier for recognizing sounds of new and learned classes. The model is first trained with a random episodic training strategy, and then its backbone is used to generate the prototypes. A dynamic relation projection module refines the prototypes to enhance their discriminability. Results on two datasets (derived from the corpora of Nsynth and FSD-MIX-CLIPS) show that the proposed method exceeds three state-of-the-art methods in average accuracy and performance dropping rate.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Few-shot classification | FSC-89 → NSynth-100 | Accuracy34.32 | 31 | |
| Audio Classification | FSC-89 | Average Accuracy (AA)35.46 | 25 | |
| Few-shot classification | FSC-89 → LS-100 | Accuracy37.66 | 22 | |
| Audio Classification | LS-100 | Accuracy (Session 0)92.35 | 13 | |
| Few-shot classification | NSynth-100 to FSC-89 (cross-dataset) | AA84.04 | 13 | |
| Few-shot classification | LS-100 to FSC-89 (cross-dataset) | AA76.23 | 13 | |
| Few-shot classification | LS-100 to NSynth-100 (cross-dataset) | Average Accuracy74.32 | 13 | |
| Class-incremental learning | NSynth 100 (test) | Accuracy (Session 0)99.96 | 13 | |
| Few-shot classification | NSynth-100 to LS-100 (cross-dataset) | AA76.54 | 13 |