Human Instance Segmentation and Tracking via Data Association and Single-stage Detector

03/31/2022
by   Lu Cheng, et al.
0

Human video instance segmentation plays an important role in computer understanding of human activities and is widely used in video processing, video surveillance, and human modeling in virtual reality. Most current VIS methods are based on Mask-RCNN framework, where the target appearance and motion information for data matching will increase computational cost and have an impact on segmentation real-time performance; on the other hand, the existing datasets for VIS focus less on all the people appearing in the video. In this paper, to solve the problems, we develop a new method for human video instance segmentation based on single-stage detector. To tracking the instance across the video, we have adopted data association strategy for matching the same instance in the video sequence, where we jointly learn target instance appearances and their affinities in a pair of video frames in an end-to-end fashion. We have also adopted the centroid sampling strategy for enhancing the embedding extraction ability of instance, which is to bias the instance position to the inside of each instance mask with heavy overlap condition. As a result, even there exists a sudden change in the character activity, the instance position will not move out of the mask, so that the problem that the same instance is represented by two different instances can be alleviated. Finally, we collect PVIS dataset by assembling several video instance segmentation datasets to fill the gap of the current lack of datasets dedicated to human video segmentation. Extensive simulations based on such dataset has been conduct. Simulation results verify the effectiveness and efficiency of the proposed work.

READ FULL TEXT

page 1

page 4

page 7

page 8

research
11/30/2020

End-to-End Video Instance Segmentation with Transformers

Video instance segmentation (VIS) is the task that requires simultaneous...
research
10/22/2021

1st Place Solution for the UVO Challenge on Video-based Open-World Segmentation 2021

In this report, we introduce our (pretty straightforard) two-step "detec...
research
06/12/2018

A Graph Transduction Game for Multi-target Tracking

Semi-supervised learning is a popular class of techniques to learn from ...
research
12/14/2020

DeepGamble: Towards unlocking real-time player intelligence using multi-layer instance segmentation and attribute detection

Annually the gaming industry spends approximately 15 billion in marketin...
research
08/29/2023

NOVIS: A Case for End-to-End Near-Online Video Instance Segmentation

Until recently, the Video Instance Segmentation (VIS) community operated...
research
06/07/2021

Contextual Guided Segmentation Framework for Semi-supervised Video Instance Segmentation

In this paper, we propose Contextual Guided Segmentation (CGS) framework...
research
03/30/2023

MobileInst: Video Instance Segmentation on the Mobile

Although recent approaches aiming for video instance segmentation have a...

Please sign up or login with your details

Forgot password? Click here to reset