Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

About

Discrete speech tokens obtained from self-supervised learning (SSL) models provide efficient data compression while maintaining strong performance, and have been widely used as intermediate representations in various tasks. However, discretization inevitably causes information loss, leading to degraded performance compared with continuous SSL features. In this work, we propose to apply soft token assignment only during downstream inference. This approach preserves the efficiency of hard discretization during training while enhancing the expressiveness of the tokens at inference. The proposed method outperforms conventional hard assignment on both ASR and speech synthesis tasks, and exhibits particularly strong generalizability to out-of-domain data. For ASR of non-native speech, it even surpasses models using continuous SSL features. Moreover, analysis of the resulting representations shows they align more accurately with phonemes compared with conventional hard assignment.

Kentaro Onda, Satoru Fukayama, Daisuke Saito, Nobuaki Minematsu• 2026

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech 100h (test-clean)
WER3.7
64
Automatic Speech RecognitionCHiME-4 (dev-real)
WER17.8
44
Automatic Speech RecognitionLibriSpeech 100h (test-other)
Word Error Rate6.3
42
Automatic Speech RecognitionERJ 1.0 (non-native)
WER38.8
16
Speech ResynthesisLJSpeech
MCD5.46
7
Voice ConversionTIMIT
PPG Distance0.808
7
Showing 6 of 6 rows

Other info

Follow for update