Cross-modal Orthogonal High-rank Augmentation for RGB-Event Transformer-trackers

07/09/2023
by   Zhiyu Zhu, et al.
0

This paper addresses the problem of cross-modal object tracking from RGB videos and event data. Rather than constructing a complex cross-modal fusion network, we explore the great potential of a pre-trained vision Transformer (ViT). Particularly, we delicately investigate plug-and-play training augmentations that encourage the ViT to bridge the vast distribution gap between the two modalities, enabling comprehensive cross-modal information interaction and thus enhancing its ability. Specifically, we propose a mask modeling strategy that randomly masks a specific modality of some tokens to enforce the interaction between tokens from different modalities interacting proactively. To mitigate network oscillations resulting from the masking strategy and further amplify its positive effect, we then theoretically propose an orthogonal high-rank loss to regularize the attention matrix. Extensive experiments demonstrate that our plug-and-play training augmentation techniques can significantly boost state-of-the-art one-stream and twostream trackers to a large extent in terms of both tracking precision and success rate. Our new perspective and findings will potentially bring insights to the field of leveraging powerful pre-trained ViTs to model cross-modal data. The code will be publicly available.

READ FULL TEXT

page 4

page 7

page 8

research
11/08/2021

Cross-Modal Object Tracking: Modality-Aware Representations and A Unified Benchmark

In many visual systems, visual tracking often bases on RGB image sequenc...
research
02/16/2023

Hierarchical Cross-modal Transformer for RGB-D Salient Object Detection

Most of existing RGB-D salient object detection (SOD) methods follow the...
research
07/31/2021

Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking

The target representation learned by convolutional neural networks plays...
research
03/01/2017

RGB-D Salient Object Detection Based on Discriminative Cross-modal Transfer Learning

In this work, we propose to utilize Convolutional Neural Networks to boo...
research
06/24/2021

A Transformer-based Cross-modal Fusion Model with Adversarial Training for VQA Challenge 2021

In this paper, inspired by the successes of visionlanguage pre-trained m...
research
06/15/2023

Cross-Modal Video to Body-joints Augmentation for Rehabilitation Exercise Quality Assessment

Exercise-based rehabilitation programs have been shown to enhance qualit...
research
12/12/2022

Cross-Modal Learning with 3D Deformable Attention for Action Recognition

An important challenge in vision-based action recognition is the embeddi...

Please sign up or login with your details

Forgot password? Click here to reset