Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer

About

Household speaker identification with few enrollment utterances is an important yet challenging problem, especially when household members share similar voice characteristics and room acoustics. A common embedding space learned from a large number of speakers is not universally applicable for the optimal identification of every speaker in a household. In this work, we first formulate household speaker identification as a few-shot open-set recognition task and then propose a novel embedding adaptation framework to adapt speaker representations from the given universal embedding space to a household-specific embedding space using a set-to-set function, yielding better household speaker identification performance. With our algorithm, Open-set Few-shot Embedding Adaptation with Transformer (openFEAT), we observe that the speaker identification equal error rate (IEER) on simulated households with 2 to 7 hard-to-discriminate speakers is reduced by 23% to 31% relative.

Kishan K C, Zhenning Tan, Long Chen, Minho Jin, Eunjung Han, Andreas Stolcke, Chul Lee• 2022

Related benchmarks

TaskDatasetResultRank
Few-shot Audio ClassificationFSC-89, NSynth-100, and LS-100 Generalizability Cross-Dataset 5-way 5-shot
Accuracy54.37
54
Few-shot classificationFSC-89 → NSynth-100
Accuracy25.64
31
Few-shot classificationFSC-89 → LS-100
Accuracy54.37
22
Few-shot Open-set Audio ClassificationDomestic Environments
Accuracy70.88
18
Few-shot classificationFSC-89 to NSynth-100
AUROC0.2386
18
Few-shot classificationFSC-89 to LS-100
AUROC56.12
9
Few-shot classificationLS-100 to NSynth-100
AUROC62.88
9
Few-shot classificationNSynth 100 to LS-100
AUROC35.56
9
Few-shot classificationNSynth-100 → LS-100
Acc38.28
9
Few-shot classificationLS-100 to FSC-89
AUROC31.69
9
Showing 10 of 18 rows

Other info

Follow for update