Explicit Visual Prompting for Universal Foreground Segmentations

05/29/2023
by   Weihuang Liu, et al.
0

Foreground segmentation is a fundamental problem in computer vision, which includes salient object detection, forgery detection, defocus blur detection, shadow detection, and camouflage object detection. Previous works have typically relied on domain-specific solutions to address accuracy and robustness issues in those applications. In this paper, we present a unified framework for a number of foreground segmentation tasks without any task-specific designs. We take inspiration from the widely-used pre-training and then prompt tuning protocols in NLP and propose a new visual prompting model, named Explicit Visual Prompting (EVP). Different from the previous visual prompting which is typically a dataset-level implicit embedding, our key insight is to enforce the tunable parameters focusing on the explicit visual content from each individual image, i.e., the features from frozen patch embeddings and high-frequency components. Our method freezes a pre-trained model and then learns task-specific knowledge using a few extra parameters. Despite introducing only a small number of tunable parameters, EVP achieves superior performance than full fine-tuning and other parameter-efficient fine-tuning methods. Experiments in fourteen datasets across five tasks show the proposed method outperforms other task-specific methods while being considerably simple. The proposed method demonstrates the scalability in different architectures, pre-trained weights, and tasks. The code is available at: https://github.com/NiFangBaAGe/Explicit-Visual-Prompt.

READ FULL TEXT

page 4

page 5

page 9

page 11

research
03/20/2023

Explicit Visual Prompting for Low-Level Structure Segmentations

We consider the generic problem of detecting low-level structures in ima...
research
09/18/2023

Parameter-Efficient Long-Tailed Recognition

The "pre-training and fine-tuning" paradigm in addressing long-tailed re...
research
07/21/2023

Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation

Parameter Efficient Tuning (PET) has gained attention for reducing the n...
research
12/19/2022

Million-scale Object Detection with Large Vision Model

Over the past few years, developing a broad, universal, and general-purp...
research
09/01/2022

Visual Prompting via Image Inpainting

How does one adapt a pre-trained visual model to novel downstream tasks ...
research
01/02/2023

Task-specific Scene Structure Representations

Understanding the informative structures of scenes is essential for low-...
research
06/13/2023

One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

We present Generalized LoRA (GLoRA), an advanced approach for universal ...

Please sign up or login with your details

Forgot password? Click here to reset