Are Multimodal Transformers Robust to Missing Modality?

04/12/2022
by   Mengmeng Ma, et al.
0

Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly preferred. Recently, Transformer models have shown great success in processing multimodal data. However, existing work has been limited to either architecture designs or pre-training strategies; whether Transformer models are naturally robust against missing-modal data has rarely been investigated. In this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of Transformers in the presence of modal-incomplete data. Unsurprising, we find Transformer models are sensitive to missing modalities while different modal fusion strategies will significantly affect the robustness. What surprised us is that the optimal fusion strategy is dataset dependent even for the same Transformer model; there does not exist a universal strategy that works in general cases. Based on these findings, we propose a principle method to improve the robustness of Transformer models by automatically searching for an optimal fusion strategy regarding input data. Experimental validations on three benchmarks support the superior performance of the proposed method.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/26/2023

Visual Prompt Flexible-Modal Face Anti-Spoofing

Recently, vision transformer based multimodal learning methods have been...
research
09/06/2022

Fusion of Satellite Images and Weather Data with Transformer Networks for Downy Mildew Disease Detection

Crop diseases significantly affect the quantity and quality of agricultu...
research
03/06/2023

Multimodal Prompting with Missing Modalities for Visual Recognition

In this paper, we tackle two challenges in multimodal learning for visua...
research
08/26/2022

TFusion: Transformer based N-to-One Multimodal Fusion Block

People perceive the world with different senses, such as sight, hearing,...
research
04/13/2023

Efficient Multimodal Fusion via Interactive Prompting

Large-scale pre-training has brought unimodal fields such as computer vi...
research
11/16/2022

Real Estate Attribute Prediction from Multiple Visual Modalities with Missing Data

The assessment and valuation of real estate requires large datasets with...
research
08/16/2022

Efficient Multimodal Transformer with Dual-Level Feature Restoration for Robust Multimodal Sentiment Analysis

With the proliferation of user-generated online videos, Multimodal Senti...

Please sign up or login with your details

Forgot password? Click here to reset