Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

About

Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. They avoid applying gradient backpropagation in large backbones, thus significantly reducing the number of trainable parameters and high memory consumption during fine-tuning. However, since they typically employ a lightweight and learnable side network, these methods inevitably introduce additional memory and time overhead during inference, which contradicts the ultimate goal of efficient transfer learning. To address the above issue, we propose a novel approach dubbed Masked Dual Path Distillation (MDPD) to accelerate inference while retaining parameter and memory efficiency in fine-tuning with fading side networks. Specifically, MDPD develops a framework that enhances the performance by mutually distilling the frozen backbones and learnable side networks in fine-tuning, and discard the side network during inference without sacrificing accuracy. Moreover, we design a novel feature-based knowledge distillation method for the encoder structure with multiple layers. Extensive experiments on distinct backbones across vision/language-only and vision-and-language tasks demonstrate that our method not only accelerates inference by at least 25.2\% while keeping parameter and memory consumption comparable, but also remarkably promotes the accuracy compared to SOTA approaches. The source code is available at https://github.com/Zhang-VKk/MDPD.

Yutong Zhang, Jiaxin Chen, Honglin Chen, Kaiqi Zheng, Shengcai Liao, Hanwen Zhong, Weixin Li, Yunhong Wang• 2026

Related benchmarks

Task	Dataset	Result
Visual Question Answering	VQA v2 (test-dev)	Overall Accuracy75.88	712
Natural Language Understanding	GLUE	SST-296.5	551
Visual Question Answering	VQA v2 (test-std)	Accuracy76.07	486
Image Classification	VTAB 1K	Overall Mean Accuracy78.3	281
Visual Grounding	RefCOCO+ (val)	Accuracy74.05	253
Visual Grounding	RefCOCO+ (testA)	Accuracy80.46	245
Visual Question Answering	GQA (test-dev)	Accuracy60.41	236
Visual Grounding	RefCOCO+ (testB)	Accuracy64.79	219
Visual Grounding	RefCOCO (val)	Accuracy83.11	172
Visual Grounding	RefCOCO (testA)	Accuracy86.77	162

Showing 10 of 20 rows

Other info

Follow for update

@wizwand_team Discord