Slot-VPS: Object-centric Representation Learning for Video Panoptic Segmentation

12/16/2021
by   Yi Zhou, et al.
0

Video Panoptic Segmentation (VPS) aims at assigning a class label to each pixel, uniquely segmenting and identifying all object instances consistently across all frames. Classic solutions usually decompose the VPS task into several sub-tasks and utilize multiple surrogates (e.g. boxes and masks, centres and offsets) to represent objects. However, this divide-and-conquer strategy requires complex post-processing in both spatial and temporal domains and is vulnerable to failures from surrogate tasks. In this paper, inspired by object-centric learning which learns compact and robust object representations, we present Slot-VPS, the first end-to-end framework for this task. We encode all panoptic entities in a video, including both foreground instances and background semantics, with a unified representation called panoptic slots. The coherent spatio-temporal object's information is retrieved and encoded into the panoptic slots by the proposed Video Panoptic Retriever, enabling it to localize, segment, differentiate, and associate objects in a unified manner. Finally, the output panoptic slots can be directly converted into the class, mask, and object ID of panoptic objects in the video. We conduct extensive ablation studies and demonstrate the effectiveness of our approach on two benchmark datasets, Cityscapes-VPS (val and test sets) and VIPER (val set), achieving new state-of-the-art performance of 63.7, 63.3 and 56.2 VPQ, respectively.

READ FULL TEXT

page 6

page 13

page 14

page 15

page 16

research
09/30/2019

LIP: Learning Instance Propagation for Video Object Segmentation

In recent years, the task of segmenting foreground objects from backgrou...
research
03/15/2023

Guided Slot Attention for Unsupervised Video Object Segmentation

Unsupervised video object segmentation aims to segment the most prominen...
research
06/29/2017

Flow-free Video Object Segmentation

Segmenting foreground object from a video is a challenging task because ...
research
03/12/2018

Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation

The problem of video object segmentation can become extremely challengin...
research
08/19/2023

Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos

Self-supervised methods have shown remarkable progress in learning high-...
research
11/23/2018

Complementary Segmentation of Primary Video Objects with Reversible Flows

Segmenting primary objects in a video is an important yet challenging pr...
research
01/11/2021

Evaluating Disentanglement of Structured Latent Representations

We design the first multi-layer disentanglement metric operating at all ...

Please sign up or login with your details

Forgot password? Click here to reset