Spatiotemporal Pyramid Network for Video Action Recognition

03/04/2019
by   Yunbo Wang, et al.
8

Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact bilinear operator into video analysis tasks. This operator enables efficient training of bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

READ FULL TEXT

page 7

page 8

research
04/22/2016

Convolutional Two-Stream Network Fusion for Video Action Recognition

Recent applications of Convolutional Neural Networks (ConvNets) for huma...
research
03/06/2019

Semantic Adversarial Network with Multi-scale Pyramid Attention for Video Classification

Two-stream architecture have shown strong performance in video classific...
research
07/22/2018

Correlation Net : spatio temporal multimodal deep learning

This paper describes a network that is able to capture spatiotemporal co...
research
07/11/2019

Two-stream Spatiotemporal Feature for Video QA Task

Understanding the content of videos is one of the core techniques for de...
research
11/13/2018

Two-stream convolutional networks for end-to-end learning of self-driving cars

We propose a methodology to extend the concept of Two-Stream Convolution...
research
07/21/2019

Attention Filtering for Multi-person Spatiotemporal Action Detection on Deep Two-Stream CNN Architectures

Action detection and recognition tasks have been the target of much focu...
research
05/20/2018

STS Classification with Dual-stream CNN

The structured time series (STS) classification problem requires the mod...

Please sign up or login with your details

Forgot password? Click here to reset