Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

About

Audio-Visual Speech Recognition (AVSR) enhances speech recognition robustness by leveraging visual cues, while real-world scenarios remain challenging due to viewpoint variation, audio distortion, and visual occlusion, which degrade modality quality and increase audio-visual asynchrony. In this paper, we propose a novel Modality-aware Multi-view Self-supervised representation framework for robust Audio-Visual Speech Recognition (M2S-AVSR). First, we introduce a multi-view representation learning encoder to learn view-invariant visual speech representations. Next, we employ a modality-aware module that explicitly models modality quality and cross-modal synchrony to perform fine-grained modality-aware fusion, enabling fine-grained visual information injection during decoding. In addition, we release AISHELL8-RealScene, a public multi-scenario, multi-view conversational audio-visual dataset recorded in real-world environments, and establish a speech recognition benchmark on it. Experiments on English and Mandarin benchmarks demonstrate the effectiveness of the proposed method under challenging conditions. On LRS3, M2S-AVSR achieves up to 29.4% relative improvement under viewpoint perturbation and visual degradation settings. Our method also achieves new state-of-the-art performance on the MISP2021-AVSR test set. On AISHELL8-RealScene, it achieves the best result in outdoor scenes. The proposed method and dataset provide useful support for future research on robust speech and multimodal tasks under realistic conditions.

Fei Su, Cancan Li, Ming Li, Juan Liu• 2026

Related benchmarks

TaskDatasetResultRank
Audio-Visual Speech RecognitionLRS-3 Babble noise at 0dB SNR (test)
WER2.02
53
Audio-Visual Speech RecognitionLRS3 Clean original (test)
WER0.65
21
Audio-Visual Speech RecognitionLRS3 Noisy w.r.t Viewpoint 5° Offset (test)
WER3.86
16
Audio-Visual Speech RecognitionLRS3 Noisy w.r.t Viewpoint 15° Offset (test)
WER4.05
16
Audio-Visual Speech RecognitionLRS3 Noisy w.r.t Missing Modality 0.3 Masking Ratio (test)
WER (%)5.77
16
Audio-Visual Speech RecognitionLRS3 Noisy w.r.t Missing Modality 0.1 Masking Ratio (test)
Word Error Rate (%)4.6
16
Audio-Visual Speech RecognitionMISP 2021 (test)
CER0.1882
13
Speech RecognitionAISHELL-8-RealScene Outdoor (test)
CER37.47
9
Speech RecognitionAISHELL-8-RealScene Indoor 5.28h (test)
CER25.35
9
Speech RecognitionAISHELL-8-RealScene Overall 11.65h (test)
CER31.41
9
Showing 10 of 10 rows

Other info

Follow for update