MixProLAP: Mixture-Induced Uncertainty Modeling for Probabilistic Language-Audio Pretraining
About
Acoustic environments often contain multiple overlapping sound events, and the same acoustic scene can be described using diverse textual expressions, making audio-text alignment inherently ambiguous. This paper proposes a probabilistic audio-language pretraining framework to model many-to-many correspondence ambiguity in audio-text alignment. Unlike conventional contrastive methods that learn deterministic point embeddings, our approach represents each modality as a distribution and learns uncertainty-aware cross-modal alignment. Rather than relying on masking-based uncertainty simulation, we mix audio-text pairs to create overlapping sounds that better reflect real acoustic mixtures and capture semantic inclusion relations among sound events. We further introduce a multi-level inclusion loss to enforce representations consistent with these relations. Experiments on audio-text retrieval benchmarks show that the proposed method outperforms deterministic baselines.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text-to-Audio Retrieval | AudioCaps (test) | Recall@125.53 | 191 | |
| Audio-to-Text Retrieval | AudioCaps (test) | R@126.85 | 80 | |
| Text-to-Audio Retrieval | Clotho V2 (test) | R@115.62 | 17 | |
| Audio-to-Text Retrieval | Clotho V2 (test) | Recall@115.6 | 17 |