Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

About

Speech large language models (Speech LLMs) typically encode speech into sequences far longer than text, creating a major efficiency bottleneck during autoregressive decoding. A common remedy is to compress the speech sequence at the adapter level to remove temporal redundancy before it enters the LLM; however, such early downsampling risks discarding fine-grained information that cannot be recovered. We propose SpeechKV, which applies a learned pooling to the KV cache of speech tokens inside the LLM. This design allows the LLM to fuse speech and text internally while directly accelerating decoding. Trained on 71K hours of speech data, SpeechKV compresses the speech to approximately text-level granularity yet maintains performance on par with or even slightly better than the uncompressed baseline, with relative gains of 6.6% on out-of-domain entity recognition and 2.3% on OpenASR, while delivering at least 1.49 times decoding speedup that scales with audio length.

Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao, Jinyu Li• 2026

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionOpenASR
Average WER5.98
100
Automatic Speech RecognitionIn-house ASR
AVG Score5.95
11
Entity recognitionIn-house ER
Average Score14.91
11
Showing 3 of 3 rows

Other info

Follow for update