Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining

About

Acoustic environments often contain multiple overlapping sound events, and the same acoustic scene can be described using diverse textual expressions, making audio-text alignment inherently ambiguous. This paper proposes a probabilistic audio-language pretraining framework to model many-to-many correspondence ambiguity in audio-text alignment. Unlike conventional contrastive methods that learn deterministic point embeddings, our approach represents each modality as a distribution and learns uncertainty-aware cross-modal alignment. Rather than relying on masking-based uncertainty simulation, we mix audio-text pairs to create overlapping sounds that better reflect real acoustic mixtures and capture semantic inclusion relations among sound events. We further introduce a multi-level inclusion loss to enforce representations consistent with these relations. Experiments on audio-text retrieval benchmarks show that the proposed method outperforms deterministic baselines.

Yu Nakagome, Jaesong Lee, Soo-Whan Chung• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Audio RetrievalAudioCaps (test)
Recall@125.53
191
Audio-to-Text RetrievalAudioCaps (test)
R@126.85
80
Text-to-Audio RetrievalClotho V2 (test)
R@115.62
17
Audio-to-Text RetrievalClotho V2 (test)
Recall@115.6
17
Showing 4 of 4 rows

Other info

Follow for update