LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training

About

Recently, inspired by the concept of sparsity, Mixture-of-Experts (MoE) models have gained increasing popularity for scaling model size while keeping the number of activated parameters constant. In this study, we thoroughly investigate the sparsity of the dense LLaMA model by constructing MoE for both the attention (i.e., Attention MoE) and MLP (i.e., MLP MoE) modules in the transformer blocks. Specifically, we investigate different expert construction methods and granularities under the same activation conditions to analyze the impact of sparsifying the model. Additionally, to comprehensively evaluate the model's capabilities across various domains (e.g., conversation, code, math) after sparsification, we apply sparsity to the instructed large language models (LLMs) and construct instructed MoE models. To counteract the performance degradation resulting from increased sparsity, we design a two-stage post-training strategy to enhance model performance. Experiments on the LLaMA3 model demonstrate the potential effectiveness of this approach for future developments of instructed MoE models. The source codes and models are available at: \url{https://github.com/OpenSparseLLMs/LLaMA-MoE-v2}.

Xiaoye Qu, Daize Dong, Xuyang Hu, Tong Zhu, Weigao Sun, Yu Cheng• 2024

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	WinoGrande	Accuracy56.1	1442
Code Generation	HumanEval (test)	--	612
Physical Interaction Question Answering	PIQA	Accuracy67.9	415
Science Question Answering	ARC Easy	Accuracy57	162
Language Understanding	MMLU 5-shot	--	153
Language Understanding	MMLU 5-shot (test)	--	149
Science Question Answering	SciQ	Normalized Accuracy88.8	137
Logical reasoning	LogiQA	Accuracy30.7	100
Instruction Following	IFEval (test)	IFEval Score36	88
Commonsense Reasoning	HellaSwag 10-shot (test)	Accuracy53.7	34

Showing 10 of 38 rows

Other info

Follow for update

@wizwand_team Discord