Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

About

Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data. Although pretrained encoders provide strong modality-specific representations, full fine-tuning can be computationally expensive, while keeping encoders frozen may limit task-specific adaptation. We propose the Tabular-Image Adapter (TI-Adapter), a modality-specific adapter-based fine-tuning framework for efficient multimodal adaptation. TI-Adapter freezes the pretrained tabular encoder and learns an adapter after the extracted tabular embedding, while adapting the image branch with embedding-level and bottleneck-level adapters instead of full fine-tuning. Experiments on 20 tabular-image datasets show that TI-Adapter achieves competitive or better predictive performance than full fine-tuning while using substantially fewer trainable parameters. Ablation studies further demonstrate the importance of adapter placement for balancing performance and practical efficiency.

Jiaqi Luo• 2026

Related benchmarks

TaskDatasetResultRank
Meme ClassificationHatefulMemes
Accuracy73.46
65
ClassificationPetFinder
Accuracy35.03
5
Classificationglaucoma-smdg
Accuracy87.66
5
Classificationhubmap-hpa
Accuracy69.2
5
ClassificationCBIS-DDSM
Accuracy66.29
5
Classificationzooscan-zooplankton
Accuracy92.01
5
Classificationjustin-instagram
Accuracy87.84
5
Regressionpainting-price
MSE (×10^7)1.54
5
Regressionamazon-bestseller
MSE0.598
5
RegressionAmazon Packages
MSE9.02
5
Showing 10 of 20 rows

Other info

Follow for update