Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Enhancing Audio Captioning with Auxiliary AudioSet Semantics

About

Automatic Audio Captioning (AAC) seeks to generate natural language descriptions of complex acoustic scenes, bridging auditory perception and language understanding. However, word-selection indeterminacy and increasing reliance on large-scale sequence-to-sequence or LLM-based models limit practical deployment. We propose a resource-efficient AAC framework that explicitly grounds caption generation in auxiliary AudioSet semantics. Frame-level acoustic representations extracted using a ConvNeXt encoder are augmented with top-$K$ predicted AudioSet keywords, providing structured contextual cues for decoding. A compact six-layer BART-style decoder conditions on this joint acoustic-semantic representation, enabling caption generation without LLM-scale decoding. The proposed design balances semantic grounding and computational efficiency within a compact architecture. Evaluations on Clotho V2 and AudioCaps confirm competitive caption quality under practical deployment constraints.

Shubham Gupta, Adarsh Arigala, Sri Rama Murty Kodukula• 2026

Related benchmarks

TaskDatasetResultRank
Audio CaptioningAudioCaps (test)
CIDEr0.78
222
Audio CaptioningClotho V2 (test)
BLEU-10.602
12
Showing 2 of 2 rows

Other info

Follow for update