Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

About

This paper presents a novel data-free and training-free compression approach for speech foundation models using channelwise clustering via k-means. More fine-grained, mixed sparsity pruning by layer-level varying number of parameter clusters is also explored. Experiments conducted on the LibriSpeech dataset suggest that when operating with pruning sparsity of 50% on HuBERT-large, consistent WER reductions of 27.73%/18.61% absolute (34.37%/21.91% relative) over the magnitude-based pruning were obtained on the test-clean and test-other subsets before fine-tuning and 0.19%/0.79% absolute (3.36%/4.62% relative) after fine-tuning with only 3 epochs. Similar WER reductions of 2.86%/5.02% absolute (59.21%/55.29% relative) were observed against magnitudebased pruning on Whisper-large-v3 at 10% sparsity, all with no significant WER increase relative to the uncompressed baseline.

Haoning Xu, Zhaoqing Li, Huimeng Wang, Youjun Chen, Chengxi Deng, Mengzhe Geng, Xunying Liu• 2026

Related benchmarks

TaskDatasetResultRank
Automatic Speech RecognitionLibriSpeech (test-other)
WER2.86
1447
Automatic Speech RecognitionLibriSpeech clean (test)
WER1.97
1410
Automatic Speech RecognitionLibriSpeech (dev-other)
WER3.11
535
Speech RecognitionLibriSpeech clean (dev)
WER0.0189
125
Showing 4 of 4 rows

Other info

Follow for update