Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

$C^3$ASD: Multi-Level Consistency-Driven Representation Learning

About

Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as background noise, occlusion, or simultaneous modality degradation. We attribute this limitation to the absence of explicit consistency constraints that promote robust, semantically aligned representations across modalities. Without such guidance, models tend to learn fragile modality-specific shortcuts that fail under corrupted conditions. We propose $C^3$ASD, a multi-level consistency-driven framework with three complementary constraints: embedding-level inter-modality consistency aligns audio-visual representations during speech; sequence-level intra-modality consistency separates speaking and non-speaking clusters via track-aware contrastive learning; and prediction-level consistency stabilizes fusion through knowledge distillation. Extensive experiments demonstrate significant improvements under diverse audio, visual and joint corruptions, while maintaining competitive performance on clean data.

Jin Hong, Jisoo Park, Junseok Kwon• 2026

Related benchmarks

TaskDatasetResultRank
Active Speaker DetectionAVA-ActiveSpeaker (val)
mAP93.8
123
Active Speaker DetectionWASD
mAP86.1
4
Active Speaker DetectionDEMAND Object Occlusion + Noise
Performance (PARK)76.1
4
Active Speaker DetectionDEMAND Pixelated Face
PARK92
4
Active Speaker DetectionAVA-ActiveSpeaker with DEMAND noise (val)
mAP (PARK)92.5
4
Active Speaker DetectionMUSAN Object Occlusion + Noise
mAP (Babble, -10 dB)66.6
4
Active Speaker DetectionMUSAN Pixelated Face
mAP (Babble, -10 dB)87.9
4
Active Speaker DetectionAVA-ActiveSpeaker MUSAN Babble noise (val)
mAP (-10 dB)88.3
4
Active Speaker DetectionAVA-ActiveSpeaker MUSAN Music noise (val)
mAP (-10 dB)87.5
4
Active Speaker DetectionAVA-ActiveSpeaker MUSAN Natural noise (val)
mAP (-10 dB)88.3
4
Showing 10 of 10 rows

Other info

Follow for update