Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Investigating End-to-End ASR Architectures for Long Form Audio Transcription

About

This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audios. We study three categories of Automatic Speech Recognition(ASR) models based on their core architecture: (1) convolutional, (2) convolutional with squeeze-and-excitation and (3) convolutional models with attention. We selected one ASR model from each category and evaluated Word Error Rate, maximum audio length and real-time factor for each model on a variety of long audio benchmarks: Earnings-21 and 22, CORAAL, and TED-LIUM3. The model from the category of self-attention with local attention and global token has the best accuracy comparing to other architectures. We also compared models with CTC and RNNT decoders and showed that CTC-based models are more robust and efficient than RNNT on long form audio.

Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind, Somshubra Majumdar, Dima Rekesh, Vahid Noroozi, Jagadeesh Balam, Boris Ginsburg• 2023

Related benchmarks

TaskDatasetResultRank
Long-form TranscriptionEarnings-22
WER19.49
30
Long-form TranscriptionEarnings-21
WER13.84
28
Automatic Speech RecognitionTED-LIUM
WER4.98
22
Showing 3 of 3 rows

Other info

Follow for update