Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SpectCount: Spectrotemporal Counting via Synthetic Signals Improves Large Audio Language Models

About

Large audio language models (LALMs) extend large language models with an audio encoder and large-scale audio data. However, the scarcity of high-quality annotated audio data remains a fundamental bottleneck for scaling. Through probing signal detectability analysis, we identify fine-grained spectrotemporal perceptual weaknesses in a foundation LALM. To address these challenges, we propose Spectrotemporal Counting (SpectCount), a data-efficient fine-tuning approach based on fully synthetic audio signals generated on-the-fly, without relying on real-world audio, annotations, or pretrained generative models. SpectCount not only resolves the observed weaknesses but also improves performance on diverse auditory benchmarks spanning sound, music, and speech, unseen during fine-tuning. These results suggest that weakness-targeted synthetic signals provide a data-efficient path toward enhanced auditory understanding capabilities in LALMs.

Seonuk Kim, Yonghyeon Jun, Ju Yeon Kang, Jimin Hong, Yoonhyeong Lee, Nam Soo Kim• 2026

Related benchmarks

TaskDatasetResultRank
Audio ReasoningMMAR
Average Accuracy56.3
82
Multiple-choice audio understandingMMAU mini (test)
Average Accuracy78.4
39
Speech UnderstandingMMSU
Accuracy63.18
35
Audio UnderstandingMMAU (test)
Speech Score69.5
31
Auditory UnderstandingAIR-Bench Foundation
Accuracy64.85
4
Showing 5 of 5 rows

Other info

Follow for update