Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

About

Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance focus on learning optimal text prompts. However, previous approaches focus on the text encoder, leaving the potential of learnable prompts within the audio encoder unexplored. In this paper, we propose a novel framework that introduces trainable prompts into the audio encoder to capture task-specific acoustic features. We demonstrate that integrating audio-side prompt learning with existing text-side approaches enhances few-shot adaptation. Through extensive experiments across 11 datasets show that integrating our method as a plug-and-play module alongside existing text prompt tuning generally leads to performance improvements. These findings suggest that explicitly modulating the audio representation space effectively complements text-only prompting approaches. The code is available at https://github.com/hyebin-c/aspl.

Hyebin Cho, Jaehyuk Jang, Changick Kim, Joon Son Chung• 2026

Related benchmarks

TaskDatasetResultRank
Audio ClassificationESC50
Top-1 Acc96.58
73
Emotion RecognitionCREMA-D
Overall Accuracy43.41
54
Emotion RecognitionRAVDESS--
46
Vocal Sound ClassificationVocalSound
Accuracy83.03
38
Instrument ClassificationBeijing Opera
Accuracy98.6
20
Acoustic Scene ClassificationTUT 2017
Accuracy83.14
15
Instrument ClassificationNS-Instruments
Top-1 Accuracy66.55
9
Music AnalysisGT-Music-Genre
Top-1 Accuracy82.17
9
Sound Event ClassificationUrban Sound
Top-1 Accuracy82.21
9
Surveillance Event ClassificationSESA
Top-1 Accuracy94.61
9
Showing 10 of 11 rows

Other info

Follow for update