Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

About

Multi-talker Automatic Speech Recognition (MTASR) faces significant challenges in accurately transcribing overlapping speech, particularly under complex high-overlap conditions. While recent Mixture-of-Experts (MoE) approaches have shown promise, they typically rely on frame-independent routing that leads to temporal myopia, and depend solely on the downstream ASR objective, which results in implicit and ungrounded representation learning. To address these limitations, we propose Holistic Speaker-Aware Guided Experts (H-SAGE) for MoE-based MTASR. Specifically, we introduce a Speaker-Aware Global Encoder to capture long-term dependencies, supervised by an auxiliary Overlap-Aware Loss that explicitly guides the model to discern acoustic states. Furthermore, we design a Holistic Gating Mechanism to arbitrate expert selection by jointly evaluating global context and local details. Experiments on LibriSpeechMix demonstrate that H-SAGE achieves consistent improvements over strong baselines, particularly in complex scenarios, validating that explicit acoustic guidance effectively enhances expert collaboration. Our code can be found at https://github.com/NKU-HLT/H-SAGE.

Yujie Guo, Jiaming Zhou, Yuhang Jia, Yang chen, Yong Qin• 2026

Related benchmarks

TaskDatasetResultRank
Speech RecognitionLibriSpeech (test)
WER0.038
81
Speech RecognitionLibriSpeech (dev)
WER3.6
26
Multi-talker Automatic Speech RecognitionLibrispeechMix 2mix (test)
WER5.7
14
Multi-talker Automatic Speech RecognitionLibrispeechMix 2mix (dev)
WER6
14
Multi-talker Speech RecognitionLibriSpeechMix LSM-3mix zero-shot generalization (dev)
WER19.7
5
Multi-talker Speech RecognitionLibriSpeechMix LSM-3mix zero-shot generalization (test)
WER19.5
5
Showing 6 of 6 rows

Other info

Follow for update