Transformer-based Multi-Modal Learning for Multi Label Remote Sensing Image Classification

06/02/2023
by   David Hoffmann, et al.
0

In this paper, we introduce a novel Synchronized Class Token Fusion (SCT Fusion) architecture in the framework of multi-modal multi-label classification (MLC) of remote sensing (RS) images. The proposed architecture leverages modality-specific attention-based transformer encoders to process varying input modalities, while exchanging information across modalities by synchronizing the special class tokens after each transformer encoder block. The synchronization involves fusing the class tokens with a trainable fusion transformation, resulting in a synchronized class token that contains information from all modalities. As the fusion transformation is trainable, it allows to reach an accurate representation of the shared features among different modalities. Experimental results show the effectiveness of the proposed architecture over single-modality architectures and an early fusion multi-modal architecture when evaluated on a multi-modal MLC dataset. The code of the proposed architecture is publicly available at https://git.tu-berlin.de/rsim/sct-fusion.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/10/2022

Multi-Modal Fusion Transformer for Visual Question Answering in Remote Sensing

With the new generation of satellite technologies, the archives of remot...
research
06/01/2023

Learning Across Decentralized Multi-Modal Remote Sensing Archives with Federated Learning

The development of federated learning (FL) methods, which aim to learn f...
research
11/06/2021

Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks

Multi-modality data is becoming readily available in remote sensing (RS)...
research
06/30/2019

Multi-Label Product Categorization Using Multi-Modal Fusion Models

In this study, we investigated multi-modal approaches using images, desc...
research
06/23/2022

Toward Clinically Assisted Colorectal Polyp Recognition via Structured Cross-modal Representation Consistency

The colorectal polyps classification is a critical clinical examination....
research
11/21/2022

TFormer: A throughout fusion transformer for multi-modal skin lesion diagnosis

Multi-modal skin lesion diagnosis (MSLD) has achieved remarkable success...
research
05/13/2021

Robust Dynamic Multi-Modal Data Fusion: A Model Uncertainty Perspective

This paper is concerned with multi-modal data fusion (MMDF) under unexpe...

Please sign up or login with your details

Forgot password? Click here to reset