Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Multimodal Transformer for Unaligned Multimodal Language Sequences

About

Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise crossmodal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.

Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov• 2019

Related benchmarks

TaskDatasetResultRank
Multimodal Sentiment AnalysisCMU-MOSEI (test)
F1 Score82.31
406
Multimodal Sentiment AnalysisCMU-MOSI (test)
F183.9
390
Multimodal Sentiment AnalysisMOSEI
MAE0.559
210
Alzheimer stage classificationADNI
Macro F176.5
200
Multimodal Sentiment AnalysisCMU-MOSI
F1 Score79.46
179
Mortality PredictionMIMIC IV
Accuracy77.45
178
Emotion Recognition in ConversationIEMOCAP (test)--
168
Mortality PredictionMIMIC IV
F1-score0.4396
154
Emotion RecognitionIEMOCAP--
151
Multimodal Sentiment AnalysisMOSI
MAE0.871
132
Showing 10 of 134 rows
...

Other info

Code

Follow for update