M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition
About
Audio-Visual Speech Recognition (AVSR) enhances speech recognition robustness by leveraging visual cues, while real-world scenarios remain challenging due to viewpoint variation, audio distortion, and visual occlusion, which degrade modality quality and increase audio-visual asynchrony. In this paper, we propose a novel Modality-aware Multi-view Self-supervised representation framework for robust Audio-Visual Speech Recognition (M2S-AVSR). First, we introduce a multi-view representation learning encoder to learn view-invariant visual speech representations. Next, we employ a modality-aware module that explicitly models modality quality and cross-modal synchrony to perform fine-grained modality-aware fusion, enabling fine-grained visual information injection during decoding. In addition, we release AISHELL8-RealScene, a public multi-scenario, multi-view conversational audio-visual dataset recorded in real-world environments, and establish a speech recognition benchmark on it. Experiments on English and Mandarin benchmarks demonstrate the effectiveness of the proposed method under challenging conditions. On LRS3, M2S-AVSR achieves up to 29.4% relative improvement under viewpoint perturbation and visual degradation settings. Our method also achieves new state-of-the-art performance on the MISP2021-AVSR test set. On AISHELL8-RealScene, it achieves the best result in outdoor scenes. The proposed method and dataset provide useful support for future research on robust speech and multimodal tasks under realistic conditions.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio-Visual Speech Recognition | LRS-3 Babble noise at 0dB SNR (test) | WER2.02 | 53 | |
| Audio-Visual Speech Recognition | LRS3 Clean original (test) | WER0.65 | 21 | |
| Audio-Visual Speech Recognition | LRS3 Noisy w.r.t Viewpoint 5° Offset (test) | WER3.86 | 16 | |
| Audio-Visual Speech Recognition | LRS3 Noisy w.r.t Viewpoint 15° Offset (test) | WER4.05 | 16 | |
| Audio-Visual Speech Recognition | LRS3 Noisy w.r.t Missing Modality 0.3 Masking Ratio (test) | WER (%)5.77 | 16 | |
| Audio-Visual Speech Recognition | LRS3 Noisy w.r.t Missing Modality 0.1 Masking Ratio (test) | Word Error Rate (%)4.6 | 16 | |
| Audio-Visual Speech Recognition | MISP 2021 (test) | CER0.1882 | 13 | |
| Speech Recognition | AISHELL-8-RealScene Outdoor (test) | CER37.47 | 9 | |
| Speech Recognition | AISHELL-8-RealScene Indoor 5.28h (test) | CER25.35 | 9 | |
| Speech Recognition | AISHELL-8-RealScene Overall 11.65h (test) | CER31.41 | 9 |