Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

EvoGM: Learning to Merge LLMs via Evolutionary Generative Optimization

About

Evolutionary model merging provides a powerful framework for the automated, training-free composition of LLMs through parameter-space search. However, existing methods predominantly rely on stochastic, hand-crafted operators that overlook the underlying performance landscape of the coefficient space. We propose Evolutionary Generative Merging (EvoGM), a framework that transcends manual heuristics by employing learnable generative modeling to optimize merging coefficients. Specifically, EvoGM features a dual-generator architecture with cycle-consistent learning to adaptively sample and refine promising merging candidates. By constructing winner-loser pairs from historical search trajectories, our framework effectively captures high-performance parameter distributions and maximizes data efficiency. This generative process is seamlessly integrated into a multi-round evolutionary pipeline, where elite merged models iteratively serve as new expert foundations. Extensive experiments across diverse benchmarks demonstrate that EvoGM significantly outperforms state-of-the-art baselines, exhibiting robust performance on both seen and unseen tasks. Code and data are available at https://github.com/JiangTao97/evogm.

Tao Jiang, Xinmeng Yu, Chenhao Yi, Yiling Wu, Yan Li, Ran Cheng, Dongmei Jiang, Jianguo Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Language UnderstandingMMLU (test)--
167
Mathematical ReasoningGSM8K (val)
Accuracy49.5
115
Multi-task Language UnderstandingMMLU (test)
Normalized Accuracy57.6
107
Multitask Language UnderstandingMMLU (val)
Accuracy64
94
Common Sense ReasoningHELLASWAG (test)
Accuracy59.4
86
Natural Language UnderstandingGLUE
Average Score (GLUE)82.4
76
Mathematical ReasoningGSM8K (test)
Accuracy (ACC)24.8
76
Commonsense ReasoningHellaSwag (val)
Accuracy66
68
Image ClassificationVision Datasets 20 tasks 1.0 (test)
Average Accuracy97.83
35
Truthfulness EvaluationTruthfulQA (test)--
30
Showing 10 of 26 rows

Other info

Follow for update