Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

5% > 100%: Flatness Preference is All You Need for Multimodal Parameter-Efficient Fine-Tuning

About

Parameter-Efficient Fine-Tuning (PEFT) methods provide a streamlined and efficient tool for adapting large models to domain-specific multimodal downstream tasks. Although these methods proved their tangible effects in practice, their principal aspects remain under-explored. Therefore we remain curious about the underlying generalization mechanisms in various PEFT methods and how they can be further enhanced. In this paper, we reveal the flatness preference widely present in various PEFTs, where a small fraction of sharp dimensions dominates the generalization of PEFT. This finding suggests an appealing possibility: we may be satisfied with a better generalization by merely attending to this small fraction of sharp dimensions instead of all of them. Furthermore, we propose Flatness Preference Optimization (FlatPO) to flatten these key sharpness dimensions, leading various PEFTs toward better generalization. Extensive experiments demonstrate the effectiveness of our findings and the proposed method. Code is available at https://github.com/Can-Lin/FlatPO.

Yifan Zhu, Can Lin, Hangjie Yuan, Zixiang Zhao, Pengfei Zhang, Tao Feng, Zhonghong Ou• 2026

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringOCRVQA
Accuracy71.4
62
Language UnderstandingGLUE and SuperGLUE fine-tuning
SST-2 Accuracy96.1
36
Visual Question AnsweringOKVQA
Accuracy53.8
34
Vision-Language Question AnsweringScienceQA
Accuracy92.4
14
Icon Question AnsweringIconQA (test)--
13
Vision-Language Question AnsweringVizWiz
Accuracy71.2
8
Vision-Language Question AnsweringIconQA txt
Accuracy93.7
8
Vision-Language Question AnsweringVQA v2
Accuracy (VQA v2)76.7
8
Showing 8 of 8 rows

Other info

Follow for update