Constraining to Generalize: Subspace Tuning for Few-shot Generalization of Audio-Language Models
About
Few-shot adaptation of pretrained Audio--Language Models (ALMs) often improves seen-class performance at the cost of unseen-class generalization, leading to the base-to-new trade-off. We attribute this failure to zero-shot drift in the text embedding space: few-shot tuning can distort inter-class structure and move adapted embeddings far from their pretrained anchors. We therefore propose Subspace Tuning (SubT), a geometry-constrained adaptation framework with two complementary controls on drift. Structured Subspace Parameterization limits structural deformation, and Residual Anchoring stabilizes adaptation around the zero-shot prior. At inference time, Subspace-aware Gating further suppresses negative transfer for weakly aligned unseen classes. Across 11 audio benchmarks, SubT delivers strong few-shot generalization while remaining efficient, operating directly on precomputed text embeddings without text-encoder backpropagation.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Classification | Beijing Opera | Base Accuracy100 | 48 | |
| Emotion Recognition | RAVDESS | Accuracy48.67 | 46 | |
| Image Classification | ImageNet Base-to-New | H Score72.56 | 44 | |
| Audio Classification | VocalSound | -- | 32 | |
| Urban Sound Classification | UrbanSound 8k | Accuracy50.98 | 29 | |
| Audio Classification | ESC50 Actions | Accuracy (Base)100 | 27 | |
| Audio Classification | RAVDESS | Base Accuracy72.85 | 27 | |
| Audio Classification | GT-Music-Genre | Base Accuracy91.09 | 27 | |
| Audio Classification | TUT 2017 | Base Accuracy93.69 | 27 | |
| Audio Classification | NS-Instruments | Base Accuracy70.66 | 27 |