Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation

08/11/2021
by   Xiaoqi Zhao, et al.
6

Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS_16, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at <https://github.com/Xiaoqi-Zhao-DLUT/Multi-Source-APS-ZVOS>.

READ FULL TEXT

page 1

page 2

page 3

page 4

page 7

page 8

research
03/18/2023

Adaptive Multi-source Predictor for Zero-shot Video Object Segmentation

Both static and moving objects usually exist in real-life videos. Most v...
research
04/08/2023

Co-attention Propagation Network for Zero-Shot Video Object Segmentation

Zero-shot video object segmentation (ZS-VOS) aims to segment foreground ...
research
03/09/2020

Motion-Attentive Transition for Zero-Shot Video Object Segmentation

In this paper, we present a novel Motion-Attentive Transition Network (M...
research
08/01/2022

BATMAN: Bilateral Attention Transformer in Motion-Appearance Neighboring Space for Video Object Segmentation

Video Object Segmentation (VOS) is fundamental to video understanding. T...
research
01/10/2023

Video Semantic Segmentation with Inter-Frame Feature Fusion and Inner-Frame Feature Refinement

Video semantic segmentation aims to generate accurate semantic maps for ...
research
10/25/2022

GlobalFlowNet: Video Stabilization using Deep Distilled Global Motion Estimates

Videos shot by laymen using hand-held cameras contain undesirable shaky ...
research
06/08/2020

Multimodal Future Localization and Emergence Prediction for Objects in Egocentric View with a Reachability Prior

In this paper, we investigate the problem of anticipating future dynamic...

Please sign up or login with your details

Forgot password? Click here to reset