The Second Place Solution for The 4th Large-scale Video Object Segmentation Challenge–Track 3: Referring Video Object Segmentation

06/24/2022
by   Leilei Cao, et al.
0

The referring video object segmentation task (RVOS) aims to segment object instances in a given video referred by a language expression in all video frames. Due to the requirement of understanding cross-modal semantics within individual instances, this task is more challenging than the traditional semi-supervised video object segmentation where the ground truth object masks in the first frame are given. With the great achievement of Transformer in object detection and object segmentation, RVOS has been made remarkable progress where ReferFormer achieved the state-of-the-art performance. In this work, based on the strong baseline framework–ReferFormer, we propose several tricks to boost further, including cyclical learning rates, semi-supervised approach, and test-time augmentation inference. The improved ReferFormer ranks 2nd place on CVPR2022 Referring Youtube-VOS Challenge.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/05/2023

ZJU ReLER Submission for EPIC-KITCHEN Challenge 2023: Semi-Supervised Video Object Segmentation

The Associating Objects with Transformers (AOT) framework has exhibited ...
research
06/02/2021

Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation

Referring video object segmentation (RVOS) aims to segment video objects...
research
09/30/2019

Towards Good Practices for Video Object Segmentation

Semi-supervised video object segmentation is an interesting yet challeng...
research
05/02/2022

Boosting Video Object Segmentation based on Scale Inconsistency

We present a refinement framework to boost the performance of pre-traine...
research
07/16/2020

Kernelized Memory Network for Video Object Segmentation

Semi-supervised video object segmentation (VOS) is a task that involves ...
research
05/02/2019

The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation

We present the 2019 DAVIS Challenge on Video Object Segmentation, the th...
research
01/03/2022

Language as Queries for Referring Video Object Segmentation

Referring video object segmentation (R-VOS) is an emerging cross-modal t...

Please sign up or login with your details

Forgot password? Click here to reset