Adaptive Temporal Encoding Network for Video Instance-level Human Parsing

08/02/2018
by   Qixian Zhou, et al.
2

Beyond the existing single-person and multiple-person human parsing tasks in static images, this paper makes the first attempt to investigate a more realistic video instance-level human parsing that simultaneously segments out each person instance and parses each instance into more fine-grained parts (e.g., head, leg, dress). We introduce a novel Adaptive Temporal Encoding Network (ATEN) that alternatively performs temporal encoding among key frames and flow-guided feature propagation from other consecutive frames between two key frames. Specifically, ATEN first incorporates a Parsing-RCNN to produce the instance-level parsing result for each key frame, which integrates both the global human parsing and instance-level human segmentation into a unified model. To balance between accuracy and efficiency, the flow-guided feature propagation is used to directly parse consecutive frames according to their identified temporal consistency with key frames. On the other hand, ATEN leverages the convolution gated recurrent units (convGRU) to exploit temporal changes over a series of key frames, which are further used to facilitate the frame-level instance-level parsing. By alternatively performing direct feature propagation between consistent frames and temporal encoding network among key frames, our ATEN achieves a good balance between frame-level accuracy and time efficiency, which is a common crucial problem in video object segmentation research. To demonstrate the superiority of our ATEN, extensive experiments are conducted on the most popular video segmentation benchmark (DAVIS) and a newly collected Video Instance-level Parsing (VIP) dataset, which is the first video instance-level human parsing dataset comprised of 404 sequences and over 20k frames with instance-level and pixel-wise annotations.

READ FULL TEXT

page 1

page 3

page 8

research
08/01/2018

Instance-level Human Parsing via Part Grouping Network

Instance-level human parsing towards real-world human analysis scenarios...
research
11/29/2016

Surveillance Video Parsing with Single Frame Supervision

Surveillance video parsing, which segments the video frames into several...
research
07/16/2020

SiamParseNet: Joint Body Parsing and Label Propagation in Infant Movement Videos

General movement assessment (GMA) of infant movement videos (IMVs) is an...
research
07/14/2022

AIParsing: Anchor-free Instance-level Human Parsing

Most state-of-the-art instance-level human parsing models adopt two-stag...
research
12/01/2016

Video Scene Parsing with Predictive Feature Learning

In this work, we address the challenging video scene parsing problem by ...
research
11/28/2021

CDGNet: Class Distribution Guided Network for Human Parsing

The objective of human parsing is to partition a human in an image into ...
research
10/02/2019

Object Parsing in Sequences Using CoordConv Gated Recurrent Networks

We present a monocular object parsing framework for consistent keypoint ...

Please sign up or login with your details

Forgot password? Click here to reset