Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition

About

Non-autoregressive (NAR) decoding generates output tokens in parallel, making speech recognition faster than autoregressive decoding, which generates them sequentially from left to right. However, the recognition performance is degraded because NAR decoding cannot resolve uncertainty by conditioning on previously generated tokens. To address this issue, we propose a novel NAR decoding framework based on minimum Bayes' risk (MBR) decoding, termed NAR-MBR decoding, that maximizes the expected utility calculated from samples drawn from the output probability of an NAR model rather than maximizing the output probability. Notably, by leveraging the nature of NAR models, multiple samples are obtained efficiently with a single forward computation. Our experiments across LibriSpeech, Switchboard, AMI, and web presentation corpus demonstrated that our NAR-MBR decoding outperformed previous NAR decoding and ran faster than AR decoding.

Hiroyuki Deguchi, Takatomo Kano, Katsuki Chousa, Marc Delcroix• 2026

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech Other
WER7
140
Automatic Speech RecognitionLibriSpeech Clean
WER3.1
124
Automatic Speech RecognitionAMI
WER18.1
46
Speech RecognitionSwitchboard
WER7.3
37
Automatic Speech RecognitionSWITCHBOARD callhm
WER14.9
14
Automatic Speech RecognitionWeb presentation corpus
WER7.3
11
Speech RecognitionWEB
Relative Speed Factor75.2
9
Speech RecognitionLibriSpeech Clean
Relative Speed (vs Beam)38.7
9
Speech RecognitionLibriSpeech Other
Relative Speed (vs Beam)32.1
9
Showing 9 of 9 rows

Other info

Follow for update