QERA: an Analytical Framework for Quantization Error Reconstruction

About

The growing number of parameters and computational demands of large language models (LLMs) present significant challenges for their efficient deployment. Recently, there is an increasing interest in quantizing weights to extremely low precision while offsetting the resulting error with low-rank, high-precision error reconstruction terms. The combination of quantization and low-rank approximation is now popular in both adapter-based, parameter-efficient fine-tuning methods such as LoftQ and low-precision inference techniques including ZeroQuant-V2. Usually, the low-rank terms are calculated via the singular value decomposition (SVD) of the weight quantization error, minimizing the Frobenius and spectral norms of the weight approximation error. Recent methods like LQ-LoRA and LQER introduced hand-crafted heuristics to minimize errors in layer outputs (activations) rather than weights, resulting improved quantization results. However, these heuristic methods lack an analytical solution to guide the design of quantization error reconstruction terms. In this paper, we revisit this problem and formulate an analytical framework, named Quantization Error Reconstruction Analysis (QERA), and offer a closed-form solution to the problem. We show QERA benefits both existing low-precision fine-tuning and inference methods -- QERA achieves a fine-tuned accuracy gain of $\Delta_{\text{acc}}$ = 6.05% of 2-bit RoBERTa-base on GLUE compared to LoftQ; and obtains $\Delta_{\text{acc}}$ = 2.97% higher post-training quantization accuracy of 4-bit Llama-3.1-70B on average than ZeroQuant-V2 and $\Delta_{\text{ppl}}$ = - 0.28 lower perplexity on WikiText2 than LQER.

Cheng Zhang, Jeffrey T. H. Wong, Can Xiao, George A. Constantinides, Yiren Zhao• 2024

Related benchmarks

Task	Dataset	Result
Language Modeling	WikiText-2 (test)	PPL4.98	2333
Language Modeling	C4	Perplexity7.88	1688
Question Answering	ARC Challenge	Accuracy47.33	906
Question Answering	ARC Easy	--	597
Question Answering	PIQA	Accuracy78	505
Sentence Completion	HellaSwag	Accuracy67.33	364
Question Answering	BoolQ	--	317
Word Prediction	LAMBADA	Accuracy73	192
Zero-shot Evaluation	Evaluation Tasks Zero-shot Aggregate	Avg. Accuracy71.9	74
Pronoun Resolution	WinoGrande	Accuracy75.33	58

Showing 10 of 14 rows

Other info

Follow for update

@wizwand_team Discord