Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Fair-GPTQ: Bias-Aware Quantization for Large Language Models

About

The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However, recent empirical studies show that, while efficient, quantization can increase the likelihood of generating biased outputs and degrade performance on fairness benchmarks. In this work, we draw new links between quantization and model fairness by adding explicit group-fairness constraints to the quantization objective and introduce Fair-GPTQ, the first quantization method explicitly designed to reduce unfairness in large language models. The added constraints guide the learning of the rounding operation toward less-biased text generation for protected groups. Specifically, we focus on stereotype generation involving occupational bias and discriminatory language spanning gender, race, and religion. Fair-GPTQ has minimal impact on performance, preserving at least 90% of baseline accuracy on zero-shot benchmarks, reduces unfairness relative to a half-precision model, and retains the memory and speed benefits of 4-bit quantization.

Irina Proskurina, Guillaume Metzler, Julien Velcin• 2025

Related benchmarks

TaskDatasetResultRank
Bias EvaluationBBQ--
175
Language ModelingWikiText
Perplexity6.28
47
Fairness evaluationCrowS-Pairs
Score67.26
34
Fairness evaluationUNQOVER
Score77.76
34
Commonsense ReasoningHellaSwag
HellaSwag Zero-shot Accuracy60.57
25
Fairness evaluationCo-occurrence tests (CooC)
CooC Score90.88
18
Efficiency and Optimization EvaluationMulti-Objective Efficiency Evaluation
Latency4.12
18
Question AnsweringAI2 Reasoning Challenge Easy
ARC-e Accuracy81.27
18
Showing 8 of 8 rows

Other info

Follow for update