Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

About

Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acoustic features. We present CARD, an encoder-free audio captioning model that removes the encoder at inference: a 13.2M projector feeds a frozen LLM with merged LoRA adapters, while the teacher used to train it is discarded. CARD distills a pretrained audio teacher (CLAP-HTSAT) into the model, but rather than injecting it into the LLM alone, it routes the teacher's representations across components: perceptual stages to the projector and semantic stages to the LLM. This placement improves CIDEr-D by +12.18 over an LLM-only distilled model on AudioCaps and by +5.21 on Clotho, reaching 55.4 against a 66.4 encoder-kept upper bound with no encoder at inference, showing that where a teacher's knowledge is placed matters as much as its presence.

Ganesh Pavan Kartikeya Bharadwaj Kolluri, Yuchen Zhang, Michael Kampouridis, Ravi Shekhar• 2026

Related benchmarks

TaskDatasetResultRank
Automated Audio CaptioningAudioCaps (evaluation)
SPIDEr35.2
17
Audio CaptioningClotho (evaluation)
CIDEr-D27.5
8
Showing 2 of 2 rows

Other info

Follow for update