Hierarchical Video Frame Sequence Representation with Deep Convolutional Graph Network

06/02/2019
by   Feng Mao, et al.
0

High accuracy video label prediction (classification) models are attributed to large scale data. These data could be frame feature sequences extracted by a pre-trained convolutional-neural-network, which promote the efficiency for creating models. Unsupervised solutions such as feature average pooling, as a simple label-independent parameter-free based method, has limited ability to represent the video. While the supervised methods, like RNN, can greatly improve the recognition accuracy. However, the video length is usually long, and there are hierarchical relationships between frames across events in the video, the performance of RNN based models are decreased. In this paper, we proposes a novel video classification method based on a deep convolutional graph neural network(DCGN). The proposed method utilize the characteristics of the hierarchical structure of the video, and performed multi-level feature extraction on the video frame sequence through the graph network, obtained a video representation re ecting the event semantics hierarchically. We test our model on YouTube-8M Large-Scale Video Understanding dataset, and the result outperforms RNN based benchmarks.

READ FULL TEXT

page 3

page 7

research
07/11/2017

Hierarchical Deep Recurrent Architecture for Video Understanding

This paper introduces the system we developed for the Youtube-8M Video U...
research
11/27/2018

A Coarse-to-fine Deep Convolutional Neural Network Framework for Frame Duplication Detection and Localization in Video Forgery

Frame duplication is to duplicate a sequence of consecutive frames and i...
research
07/04/2017

Aggregating Frame-level Features for Large-Scale Video Classification

This paper introduces the system we developed for the Google Cloud & You...
research
08/27/2018

Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions

We propose a novel attentive sequence to sequence translator (ASST) for ...
research
11/14/2014

A Discriminative CNN Video Representation for Event Detection

In this paper, we propose a discriminative video representation for even...
research
12/02/2017

Lecture video indexing using boosted margin maximizing neural networks

This paper presents a novel approach for lecture video indexing using a ...
research
02/25/2015

Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks

In this paper, we study the challenging problem of categorizing videos a...

Please sign up or login with your details

Forgot password? Click here to reset