USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
About
Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (SSL) has yielded strong domain-specific encoders like speech or music experts, multi-domain approaches like USAD and SPEAR remain limited in coverage and evaluation. Recent studies also suggest supervised encoders align better with audio LLMs. We present USAD 2.0, a universal encoder integrating knowledge from both SSL and supervised foundation models. USAD 2.0 introduces domain-aware distillation to address teacher mismatch, extends coverage to the music domain, and adds second-stage supervised distillation for downstream use. We further scale the model to one billion parameters via depth scaling. Experiments show USAD 2.0 achieves strong or state-of-the-art performance across probing and LLM-based evaluations.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Classification | ESC-50 | Accuracy96.8 | 461 | |
| Audio Representation Evaluation | HEAR (Holistic Evaluation of Audio Representations) | HEAR Average84.4 | 59 | |
| Speech Processing | SUPERB | KWS Acc0.977 | 52 | |
| Audio Event Tagging | AudioSet (AS-20K) | mAP40.9 | 39 | |
| Audio Classification | XARES-LLM Track A | Track A Score78.3 | 12 | |
| Audio Understanding | XARES-LLM Track B | Track B Score62.4 | 12 | |
| Music Probing | MARBLE | Average Score75.8 | 12 |