Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Whisfusion: Parallel ASR Decoding with Masked Diffusion

About

Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length. A natural alternative, CTC-style non-autoregressive (NAR) systems avoid this bottleneck but their conditional independence assumption sacrifices transcript-level generative modeling. Masked diffusion language models (e.g., LLaDA, MDLM) offer a competitive NAR text-generation approach. We ask whether such models can bring NAR ASR into the accuracy regime of strong AR ASR systems while removing the left-to-right bottleneck. We propose Whisfusion, which trains a dedicated masked diffusion decoder from scratch on top of frozen Whisper-large-v3 audio embeddings, denoising masked transcripts in just a few steps. We train on ~68k hours of 11-language speech with high-mask specialization to align training with the fully masked starting point of inference, and decode via Parallel Diffusion Decoding. Whisfusion surpasses Whisper-large-v3 on group-average accuracy across English, European, and CJK benchmarks, while running 4-5x faster, additionally surpassing Whisper-turbo in both accuracy and throughput. It reaches accuracy competitive with Canary and Qwen3-ASR while running 3-7x faster. These results establish masked diffusion as a Pareto-competitive non-autoregressive paradigm for high-throughput multilingual transcription. Code and model weights are available at https://github.com/taeyoun811/Whisfusion.

Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim• 2025

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech (test-other)
WER17
1447
Automatic Speech RecognitionLibriSpeech clean (test)
WER8.3
1410
Automatic Speech RecognitionLibrispeech (test-clean)
WER8.3
170
Automatic Speech RecognitionEarnings-22
WER9.64
39
Automatic Speech RecognitionLibriSpeech (LS) clean
WER1.67
21
Automatic Speech RecognitionCV-zh
CER14.49
15
Automatic Speech RecognitionAISHELL
CER4.93
12
Automatic Speech RecognitionLibriSpeech LS-other
WER3.53
10
Automatic Speech RecognitionVoxPopuli English
WER6.34
10
Automatic Speech RecognitionCommonVoice CV-en
WER9.15
10
Showing 10 of 18 rows

Other info

Follow for update