Video-ception Network: Towards Multi-Scale Efficient Asymmetric Spatial-Temporal Interactions

07/22/2020
by   Yuan Tian, et al.
4

Previous video modeling methods leverage the cubic 3D convolution filters or its decomposed variants to exploit the motion cues for precise action recognition, which tend to be performed on the video features along the temporal and spatial axes symmetrically. This brings the hypothesis implicitly that the actions are recognized from the cubic voxel level and neglects the essential spatial-temporal shape diversity across different actions. In this paper, we propose a novel video representing method that fuses the features spatially and temporally in an asymmetric way to model action atomics spanning multi-scale spatial-temporal scales. To permit the feature fusion procedure efficiently and effectively, we also design the optimized feature interaction layer, which covers most feature fusion techniques as special case of it, e.g., channel shuffling and channel concatenating. We instantiate our method as a plug-and-play block, termed Multi-Scale Efficient Asymmetric Spatial-Temporal Block. Our method can easily adapt the traditional 2D CNNs to the video understanding tasks such as action recognition. We verify our method on several most recent large-scale video datasets requiring strong temporal reasoning or appearance discriminating, e.g., Something-to-Something v1, Kinetics and Diving48, demonstrate the new state-of-the-art results without bells and whistles.

READ FULL TEXT

page 1

page 10

research
09/28/2019

Grouped Spatial-Temporal Aggregation for Efficient Action Recognition

Temporal reasoning is an important aspect of video analysis. 3D CNN show...
research
12/04/2017

Robust 3D Action Recognition through Sampling Local Appearances and Global Distributions

3D action recognition has broad applications in human-computer interacti...
research
04/03/2018

Multi-Scale Spatially-Asymmetric Recalibration for Image Classification

Convolution is spatially-symmetric, i.e., the visual features are indepe...
research
06/17/2023

Multi-scale Spatial-temporal Interaction Network for Video Anomaly Detection

Video anomaly detection (VAD) is an essential yet challenge task in sign...
research
05/19/2018

DenseImage Network: Video Spatial-Temporal Evolution Encoding and Understanding

Many of the leading approaches for video understanding are data-hungry a...
research
09/22/2022

FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video Classification

Unmanned aerial vehicles (UAVs) are now widely applied to data acquisiti...
research
06/09/2022

Spatial-temporal Concept based Explanation of 3D ConvNets

Recent studies have achieved outstanding success in explaining 2D image ...

Please sign up or login with your details

Forgot password? Click here to reset