Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Constraining to Generalize: Subspace Tuning for Few-shot Generalization of Audio-Language Models

About

Few-shot adaptation of pretrained Audio--Language Models (ALMs) often improves seen-class performance at the cost of unseen-class generalization, leading to the base-to-new trade-off. We attribute this failure to zero-shot drift in the text embedding space: few-shot tuning can distort inter-class structure and move adapted embeddings far from their pretrained anchors. We therefore propose Subspace Tuning (SubT), a geometry-constrained adaptation framework with two complementary controls on drift. Structured Subspace Parameterization limits structural deformation, and Residual Anchoring stabilizes adaptation around the zero-shot prior. At inference time, Subspace-aware Gating further suppresses negative transfer for weakly aligned unseen classes. Across 11 audio benchmarks, SubT delivers strong few-shot generalization while remaining efficient, operating directly on precomputed text embeddings without text-encoder backpropagation.

Jaehyuk Jang, Kangwook Ko, Wonjun Lee, Changick Kim• 2026

Related benchmarks

TaskDatasetResultRank
Audio ClassificationBeijing Opera
Base Accuracy100
48
Emotion RecognitionRAVDESS
Accuracy48.67
46
Image ClassificationImageNet Base-to-New
H Score72.56
44
Audio ClassificationVocalSound--
32
Urban Sound ClassificationUrbanSound 8k
Accuracy50.98
29
Audio ClassificationESC50 Actions
Accuracy (Base)100
27
Audio ClassificationRAVDESS
Base Accuracy72.85
27
Audio ClassificationGT-Music-Genre
Base Accuracy91.09
27
Audio ClassificationTUT 2017
Base Accuracy93.69
27
Audio ClassificationNS-Instruments
Base Accuracy70.66
27
Showing 10 of 18 rows

Other info

Follow for update