Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

About

Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (SSL) has yielded strong domain-specific encoders like speech or music experts, multi-domain approaches like USAD and SPEAR remain limited in coverage and evaluation. Recent studies also suggest supervised encoders align better with audio LLMs. We present USAD 2.0, a universal encoder integrating knowledge from both SSL and supervised foundation models. USAD 2.0 introduces domain-aware distillation to address teacher mismatch, extends coverage to the music domain, and adds second-stage supervised distillation for downstream use. We further scale the model to one billion parameters via depth scaling. Experiments show USAD 2.0 achieves strong or state-of-the-art performance across probing and LLM-based evaluations.

Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati, Mrudula Athi, Anton Ratnarajah, Amit Chhetri, James Glass• 2026

Related benchmarks

TaskDatasetResultRank
Audio ClassificationESC-50
Accuracy96.8
461
Audio Representation EvaluationHEAR (Holistic Evaluation of Audio Representations)
HEAR Average84.4
59
Speech ProcessingSUPERB
KWS Acc0.977
52
Audio Event TaggingAudioSet (AS-20K)
mAP40.9
39
Audio ClassificationXARES-LLM Track A
Track A Score78.3
12
Audio UnderstandingXARES-LLM Track B
Track B Score62.4
12
Music ProbingMARBLE
Average Score75.8
12
Showing 7 of 7 rows

Other info

Follow for update