Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

About

Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.

Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang• 2025

Related benchmarks

TaskDatasetResultRank
Language ModelingWikiText-2
Perplexity (PPL)5.36
2862
Physical Commonsense ReasoningPIQA
Accuracy68.39
724
Language ModelingWikiText
Word Perplexity14.31
331
Word PredictionLAMBADA
Accuracy36.81
222
ReasoningWinoGrande (WG)
Accuracy64.56
172
Question AnsweringCommonsenseQA
Accuracy23.91
172
ReasoningPIQA
Accuracy71.27
168
ReasoningARC-C
Accuracy (ARC-c)36.52
113
Commonsense ReasoningWinoGrande--
94
ReasoningARC-E
First-Token Accuracy70.75
28
Showing 10 of 19 rows

Other info

Follow for update