Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

About

Fine-tuning Multimodal Large Language Models (MLLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training dynamics across modalities often lead to suboptimal accuracy due to negative interference, a challenge typically addressed with inefficient heuristic methods such as manually tuning separate learning rates. To overcome this, we introduce MARS (Multimodal Adaptive Rank Search), an approach to discover optimal rank pairs that balance training dynamics while maximizing performance. Our key innovation, a proposed framework of dual scaling laws, enables this search: one law models module-specific convergence time to prune the search space to candidates with aligned dynamics, while the other predicts final task performance to select the optimal pair from the pruned set. By re-purposing the LoRA rank as a controller for modality-specific convergence speed, MARS outperforms baseline methods and provides a robust, automated strategy for optimizing MLLM fine-tuning.

Minkyoung Cho, Insu Jang, Shuowei Jin, Zesen Zhao, Adityan Jothi, Ethem F. Can, Min-Hung Chen, Z. Morley Mao• 2026

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE--
1455
Diagram UnderstandingAI2D
Accuracy49.9
247
Visual ReasoningGQA
Accuracy55.7
93
Multimodal Question AnsweringScienceQA (test)
Accuracy79.64
65
Image CaptioningTextCaps (val)--
51
Complex ReasoningLLaVA Bench (val)
Perplexity2.1875
44
General perception and cognitionMME SUM
Cognition + Perception Sum1.70e+3
3
Visual Perception and ReasoningMMStar avg
Average Accuracy39.4
3
Showing 8 of 8 rows

Other info

Follow for update