Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Are Multimodal Transformers Robust to Missing Modality?

About

Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly preferred. Recently, Transformer models have shown great success in processing multimodal data. However, existing work has been limited to either architecture designs or pre-training strategies; whether Transformer models are naturally robust against missing-modal data has rarely been investigated. In this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of Transformers in the presence of modal-incomplete data. Unsurprising, we find Transformer models are sensitive to missing modalities while different modal fusion strategies will significantly affect the robustness. What surprised us is that the optimal fusion strategy is dataset dependent even for the same Transformer model; there does not exist a universal strategy that works in general cases. Based on these findings, we propose a principle method to improve the robustness of Transformer models by automatically searching for an optimal fusion strategy regarding input data. Experimental validations on three benchmarks support the superior performance of the proposed method.

Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, Xi Peng• 2022

Related benchmarks

TaskDatasetResultRank
Image ClassificationFood-101 (test)
Top-1 Acc77.5
89
Multimodal Multilabel ClassificationMM-IMDB (test)
Macro F146.6
87
Readmission predictionMIMIC IV
AUC-ROC0.6901
70
Hateful Meme DetectionHateful Memes (test)
AUROC0.612
67
Multimodal ClassificationMST Missing Modalities
Accuracy99.96
28
Mortality PredictioneICU
AUC-ROC0.8882
22
Multimodal ClassificationPolyMNIST Missing Rate η=0.6
Accuracy98.43
16
Multimodal ClassificationPolyMNIST Missing Rate η=0.8
Accuracy91.14
16
Multimodal ClassificationPolyMNIST Missing Rate η=0
Accuracy99.97
16
Multimodal ClassificationMST Missing Modalities {S,T}
Accuracy0.986
14
Showing 10 of 24 rows

Other info

Follow for update