Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means

About

In recent years, there has been growing interest in representing speech with discrete tokens, which serve as pseudo-text for speech language models (speechLMs) and as efficient intermediate representations for downstream tasks. These tokens are typically categorized as acoustic and phonetic tokens: the former holds detailed acoustic information for reconstruction while the latter mainly captures linguistic content. In human speech communication, however, unnecessary acoustic details such as speaker information are abstracted, while both linguistic and prosodic information are utilized for speech comprehension and production. Given this, neither type of token seems an ideal representation for tasks sensitive to prosody, such as speechLMs. In this study, we propose the Phonological Tokenizer, a method that fine-tunes phonetic tokens via differentiable k-means with a multi-task objective of ASR and speech resynthesis. Experimental validation on diverse tasks confirms that our tokens retain phonological (both linguistic and prosodic) information while appropriately discarding speaker identity.

Kentaro Onda, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe• 2026

Related benchmarks

TaskDatasetResultRank
Speaker IdentificationVoxCeleb1
Accuracy29.5
58
Automatic Speech RecognitionLibriSpeech 100h (test-clean)
WER4.6
32
Automatic Speech RecognitionLibriSpeech 100h (test-other)
Word Error Rate8.5
10
Emotion RecognitionRAVDESS (speaker-independent)
Accuracy51.7
6
Sentiment and speaker consistency assessmentSALMon
Sentiment Accuracy67.5
6
Speech continuation quality assessmentLibriLight Speech Continuation
GenPPL5.6
6
Voice ConversionTIMIT OOD
F0 Correlation0.456
6
Voice ConversionExpresso OOD
F0 Correlation0.538
6
Lexical and syntactic knowledge assessmentZero Resource Speech Challenge
sWUGGY67
6
Speech ReconstructionLJSpeech ID
MCD4.99
6
Showing 10 of 10 rows

Other info

Follow for update