Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

About

Strong speech-to-text (S2T) LLMs already provide robust speech perception and text reasoning, but adding speech-to-speech (S2S) output is challenging: fine-tuning the backbone can degrade the original S2T performance, while attaching a downstream talker reintroduces a serial text-to-speech bottleneck. We present PRIME-Speech, a frozen-backbone S2S conversion framework that trains only speech-generation modules. PRIME-Speech synchronizes a causal audio post-decoder with intermediate hidden states of the frozen backbone, so codec tokens are generated from the model's evolving reasoning trajectory rather than from completed text chunks. The post-decoder uses mixed hidden-state, text, and audio-history conditioning, and a training-time packing strategy with turn-level audio KV-cache and position reset stabilizes multi-turn spoken interaction without additional multi-turn S2S training data. Multi-token prediction further reduces the effective codec prediction rate and improves first-audio latency without modifying the reasoning path. Across speech translation, spoken QA, speech understanding, and multi-turn dialogue, PRIME-Speech preserves the S2T behavior of the frozen backbone while producing accurate, low-WER spoken responses.

Yuxuan Hu, Heng Lu, Ruchao Fan, Yao Qian, Xiaofei Wang, Jian Xue, Heming Wang, Shuohang Wang, Young Jin Kim, Yelong Shen, Jinyu Li• 2026

Related benchmarks

TaskDatasetResultRank
Spoken Question AnsweringUltraEval-Audio LLaMA-QA
S2T Score79
9
Speech UnderstandingBigBench Audio
S2T Score66.2
9
Spoken Question AnsweringUltraEval-Audio TriviaQA
S2T Score46.98
9
Spoken Question AnsweringUltraEval-Audio WebQ
S2T Score (Task Correctness)42.04
9
Conversational Speech Question-AnsweringMulti-turn
S2T Score80.45
8
Speech Understanding and FluencyVocalBench
Knowledge Score68.9
8
Speech-to-speech translationCoVoST-2 X2EN
S2T BLEU Score41.29
7
Speech-to-speech translationFLEURS X2EN
S2T BLEU Score31.4
7
Showing 8 of 8 rows

Other info

Follow for update