Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach

About

Recent progress in Spoken Language Modeling has shown that learning language directly from speech is feasible. Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. Modeling directly from speech opens up the path to more natural and expressive systems. On the other hand, speech-only systems require up to three orders of magnitude more data to catch up to their text-based counterparts in terms of their semantic abilities. We show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.

Maxime Poli, Emmanuel Chemla, Emmanuel Dupoux• 2024

Related benchmarks

TaskDatasetResultRank
Phone recognitionTIMIT (test)--
23
Phone TranscriptionISLE (test)
WPFER5.4
9
Phone TranscriptionEpaDB (test)
WPFER8.2
9
Phone TranscriptionPSST (test)
WPFER20
9
Phone TranscriptionSpeech Ocean (test)
WPFER12.8
9
Phone TranscriptionAggregate (TIMIT, EpaDB, PSST, Speech Ocean, ISLE) (test)
Average WPFER10.6
9
Showing 6 of 6 rows

Other info

Follow for update