QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

About

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved, establishing robust correspondences between audio and visual contents poses unique challenges due to (1) complex entanglement across sound sources and (2) frequent changes in the occurrence of distinct sound events. Assuming sound events occur independently, the multi-source semantic space can be represented as the Cartesian product of single-source sub-spaces. We are motivated to decompose the multi-source audio semantics into single-source semantics for more effective interactions with visual content. We propose a semantic decomposition method based on product quantization, where the multi-source semantics can be decomposed and represented by several disentangled and noise-suppressed single-source semantics. Furthermore, we introduce a global-to-local quantization mechanism, which distills knowledge from stable global (clip-level) features into local (frame-level) ones, to handle frequent changes in audio semantics. Extensive experiments demonstrate that our semantically decomposed audio representation significantly improves AVS performance, e.g., +21.2% mIoU on the challenging AVS-Semantic benchmark with ResNet50 backbone. https://github.com/lxa9867/QSD.

Xiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng, Rita Singh, Yan Lu, Bhiksha Raj• 2023

Related benchmarks

Task	Dataset	Result
Audio-Visual Segmentation	AVSBench-Object MS3 (test)	Jaccard Index (J)61.9	21
Audio-Visual Segmentation	AVSBench AVS-Objects-MS3	J & F Score64	21
Audio-Visual Segmentation	AVSBench AVS-Objects-S4	J&F Score83.9	21
Audio-Visual Segmentation	AVSBench Object S4 (test)	Jaccard Index (J)79.5	21
Audio-Visual Segmentation	AVS-Object MS3	J&Fm Combined Score64	19
Audio-Visual Segmentation	AVS-Object S4	J&Fm83.9	19
Audio-Visual Segmentation	AVS-Object-Single	J&F Score83.9	13
Audio-Visual Segmentation	AVS-Object-Multi	J&F Score64	13
Audio-Visual Segmentation	AVS-Semantic	Jaccard Index (J)53.4	12
Audio-Visual Semantic Segmentation	AVS-Semantic	mIoU53.4	7

Showing 10 of 10 rows

Other info

Code

Follow for update

@wizwand_team Discord