Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CAAD: Contrastive Audio-Aware Distillation for Efficient Speech Language Models

About

Speech Language Models achieve reasoning capabilities, but are often hindered by massive parameter counts and a tendency to prioritize linguistic priors over acoustic features. While contrastive decoding enhances grounding by contrasting audio-aware and text-only logits, it increases inference latency. We propose Contrastive Audio-Aware Distillation (CAAD), a framework that internalizes the teacher's contrastive reasoning into the student model's weights. To overcome the high computational training overhead in the dual-path token-by-token contrastive distillation process, we introduce a synchronized teacher-forcing strategy. Anchored by unified Pseudo-Ground Truths, this mechanism enables simultaneous full-sequence generation of the teacher's contrastive distributions, allowing student to distill the audio-aware signal efficiently. Overall, CAAD yields a ~8% relative gain over standard knowledge distillation on Dynamic-SUPERB and successfully reduces linguistic bias in MCR-BENCH.

Chun-Wei Chen, Tzu-Quan Lin, Ke-Han Lu, Wei-Ping Huang, Hung-Yi Lee• 2026

Related benchmarks

TaskDatasetResultRank
Conflict ResolutionMCR-Bench
Accuracy (Neutral)45.9
6
Speech Language ModelingDynamic-SUPERB
Content Score73.86
6
Showing 2 of 2 rows

Other info

Follow for update