Iterative Pseudo-Labeling for Speech Recognition

About

Pseudo-labeling has recently shown promise in end-to-end automatic speech recognition (ASR). We study Iterative Pseudo-Labeling (IPL), a semi-supervised algorithm which efficiently performs multiple iterations of pseudo-labeling on unlabeled data as the acoustic model evolves. In particular, IPL fine-tunes an existing model at each iteration using both labeled data and a subset of unlabeled data. We study the main components of IPL: decoding with a language model and data augmentation. We then demonstrate the effectiveness of IPL by achieving state-of-the-art word-error rate on the Librispeech test sets in both standard and low-resource setting. We also study the effect of language models trained on different corpora to show IPL can effectively utilize additional text. Finally, we release a new large in-domain text corpus which does not overlap with the Librispeech training transcriptions to foster research in low-resource, semi-supervised ASR

Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, Ronan Collobert• 2020

Related benchmarks

Task	Dataset	Result
Automatic Speech Recognition	LibriSpeech (test-other)	WER4.01	1447
Automatic Speech Recognition	LibriSpeech clean (test)	WER2.1	1410
Automatic Speech Recognition	LibriSpeech (dev-other)	WER3.26	535
Automatic Speech Recognition	LibriSpeech (dev-clean)	WER (%)1.85	376
Automatic Speech Recognition	Spgispeech (test)	WER2.57	19
Automatic Speech Recognition	English DementiaBank Pitt (Eval)	WER (Paraphrase)22.49	19
Automatic Speech Recognition	Cantonese JCCOCC MoCA (dev)	CER33.69	19
Automatic Speech Recognition	Cantonese JCCOCC MoCA (eval)	CER30.46	19
Automatic Speech Recognition	Cantonese JCCOCC MoCA All	CER32.06	19
Automatic Speech Recognition	DementiaBank Pitt (dev)	WER (Parenthetical)31.71	9

Showing 10 of 18 rows

Other info

Follow for update

@wizwand_team Discord