Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

About

Large Language Models (LLMs) have demonstrated exceptional proficiency in language-related tasks, but their deployment poses significant challenges due to substantial memory and storage requirements. Weight-only quantization has emerged as a promising solution, significantly reducing memory and storage needs without sacrificing too much performance. In this study, we introduce SignRound, a method that leverages signed gradient descent (SignSGD) to optimize rounding values and weight clipping in just 200 steps. SignRound integrates the advantages of Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ), delivering exceptional results across 2 to 4 bits while minimizing tuning costs and avoiding additional inference overhead. For example, SignRound achieved absolute average accuracy improvements ranging from 6.91% to 33.22% at 2bits, as measured by the average zero-shot accuracy across 11 tasks. It also demonstrates strong generalization in recent models, achieving near-lossless 4-bit quantization in most scenarios. The source code is publicly available at https://github.com/intel/auto-round.

Wenhua Cheng, Weiwei Zhang, Haihao Shen, Yiyang Cai, Xin He, Kaokao Lv, Yi Liu• 2023

Related benchmarks

TaskDatasetResultRank
Language ModelingWiki2
PPL9.62
382
Language ModelingLAMBADA
Perplexity3.89
254
Zero-shot EvaluationPIQA, WinoGrande, HellaSwag, ARC (Easy and Challenge), LAMBADA (test)
Average Accuracy67.7
102
General Language UnderstandingLAMBADA, ARC-C, HellaSwag, BoolQ, MMLU, GSM8K Average
Accuracy72.25
56
Large Language Model Evaluation10 tasks average
Avg Accuracy69.01
50
LLM EvaluationQwen3-1.7B Evaluation Suite (avg)
Average Performance58.31
38
Large Language Model EvaluationQwen3-0.6B Average (test)
Average Performance45.75
38
Language ModelingMistral-7B--
24
Zero-shot EvaluationMobileLLM Evaluation Suite zero-shot
ARC-e26.39
23
Commonsense ReasoningCommonsense Reasoning LLaMA2-7B
Average Accuracy63.72
18
Showing 10 of 15 rows

Other info

Follow for update