Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization

About

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approaches aim to isolate semantic information from low-level acoustics to better align with language models (LMs). In particular, previous methods use self-supervised learning (SSL) teachers such as HuBERT to extract semantic representations, which are then distilled into a semantic quantizer to suppress acoustic redundancy as well as capture content-related latent structures. However, these tokenizers often operate at relatively high frame rates, producing token sequences significantly longer than their textual counterparts and hindering seamless integration with pretrained LMs. Although recent methods attempt to reduce the token rate by applying uniform average pooling to SSL features, this can over-smooth content-bearing regions and dilute the structural information, thereby potentially limiting the LM alignment. To address this, we propose LM-SPT, an LM-aligned speech tokenization method based on semantic speech-resynthesis distillation. Instead of directly matching teacher and student features via pooling, LM-SPT resynthesizes speech from semantic tokens only and minimizes the discrepancy between representations extracted from the original and resynthesized waveforms using a frozen, LM-aligned speech encoder. This indirect supervision avoids rigid temporal alignment and encourages dedicated semantic units that are more semantically aligned with LMs under reduced frame rates. Experimental results show that the proposed LM-SPT consistently outperforms previous semantic-enhanced speech tokenizers when applied to SLMs for the tasks of automatic speech recognition and text-to-speech, even without compromising the speech reconstruction fidelity at the codec level.

Daejin Jo, Jeeyoung Yun, Byungseok Roh, Sungwoong Kim• 2025

Related benchmarks

TaskDatasetResultRank
Audio Tokenization and Compression EvaluationCodec-SUPERB
SDR5.54
8
Text-to-SpeechLibriSpeech
NMOS3.89
4
Text-to-SpeechKSponSpeech
NMOS3.87
4
ReconstructionLibriSpeech and KSponSpeech
MUSHRA Score89.96
3
Showing 4 of 4 rows

Other info

Follow for update