Unifying Global-Local Representations in Salient Object Detection with Transformer

08/05/2021
by   Sucheng Ren, et al.
0

The fully convolutional network (FCN) has dominated salient object detection for a long period. However, the locality of CNN requires the model deep enough to have a global receptive field and such a deep model always leads to the loss of local details. In this paper, we introduce a new attention-based encoder, vision transformer, into salient object detection to ensure the globalization of the representations from shallow to deep layers. With the global view in very shallow layers, the transformer encoder preserves more local representations to recover the spatial details in final saliency maps. Besides, as each layer can capture a global view of its previous layer, adjacent layers can implicitly maximize the representation differences and minimize the redundant features, making that every output feature of transformer layers contributes uniquely for final prediction. To decode features from the transformer, we propose a simple yet effective deeply-transformed decoder. The decoder densely decodes and upsamples the transformer features, generating the final saliency map with less noise injection. Experimental results demonstrate that our method significantly outperforms other FCN-based and transformer-based methods in five benchmarks by a large margin, with an average of 12.17 improvement in terms of Mean Absolute Error (MAE). Code will be available at https://github.com/OliverRensu/GLSTR.

READ FULL TEXT

page 1

page 2

page 3

page 7

page 8

research
08/17/2021

Boosting Salient Object Detection with Transformer-based Asymmetric Bilateral U-Net

Existing salient object detection (SOD) methods mainly rely on CNN-based...
research
12/08/2019

SaLite : A light-weight model for salient object detection

Salient object detection is a prevalent computer vision task that has ap...
research
09/15/2023

UniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection

Video saliency prediction and detection are thriving research domains th...
research
08/20/2020

Co-Saliency Detection with Co-Attention Fully Convolutional Network

Co-saliency detection aims to detect common salient objects from a group...
research
05/24/2023

DC-Net: Divide-and-Conquer for Salient Object Detection

In this paper, we introduce Divide-and-Conquer into the salient object d...
research
07/18/2021

AS-MLP: An Axial Shifted MLP Architecture for Vision

An Axial Shifted MLP architecture (AS-MLP) is proposed in this paper. Di...
research
03/26/2023

Feature Shrinkage Pyramid for Camouflaged Object Detection with Transformers

Vision transformers have recently shown strong global context modeling c...

Please sign up or login with your details

Forgot password? Click here to reset