Training-Free Activation Sparsity in Large Language Models

About

Activation sparsity can enable practical inference speedups in large language models (LLMs) by reducing the compute and memory-movement required for matrix multiplications during the forward pass. However, existing methods face limitations that inhibit widespread adoption. Some approaches are tailored towards older models with ReLU-based sparsity, while others require extensive continued pre-training on up to hundreds of billions of tokens. This paper describes TEAL, a simple training-free method that applies magnitude-based activation sparsity to hidden states throughout the entire model. TEAL achieves 40-50% model-wide sparsity with minimal performance degradation across Llama-2, Llama-3, and Mistral families, with sizes varying from 7B to 70B. We improve existing sparse kernels and demonstrate wall-clock decoding speed-ups of up to 1.53$\times$ and 1.8$\times$ at 40% and 50% model-wide sparsity. TEAL is compatible with weight quantization, enabling further efficiency gains.

James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun• 2024

Related benchmarks

Task	Dataset	Result
Language Modeling	WikiText2	Perplexity5.15	3785
Medical Question Answering	MedMCQA	Accuracy52.95	521
Long-context Language Understanding	LongBench	--	294
General Reasoning	MMLU	MMLU Accuracy76.63	180
Question Answering	CommonsenseQA	Accuracy74.77	150
Code	HumanEval	HumanEval Accuracy46.95	118
Long-context Understanding	LongBench	Overall Average Score30.54	115
Question Answering	TruthfulQA	Accuracy57.08	73
Language Modeling	Wikitext (test)	Perplexity5.52	66
Commonsense Reasoning	Commonsense Reasoning	Accuracy74.12	57

Showing 10 of 17 rows

Other info

Follow for update

@wizwand_team Discord