MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection

12/07/2021
by   Rui Dai, et al.
5

Action detection is an essential and challenging task, especially for densely labelled datasets of untrimmed videos. The temporal relation is complex in those datasets, including challenges like composite action, and co-occurring action. For detecting actions in those complex videos, efficiently capturing both short-term and long-term temporal information in the video is critical. To this end, we propose a novel ConvTransformer network for action detection. This network comprises three main components: (1) Temporal Encoder module extensively explores global and local temporal relations at multiple temporal resolutions. (2) Temporal Scale Mixer module effectively fuses the multi-scale features to have a unified feature representation. (3) Classification module is used to learn the instance center-relative position and predict the frame-level classification scores. The extensive experiments on multiple datasets, including Charades, TSU and MultiTHUMOS, confirm the effectiveness of our proposed method. Our network outperforms the state-of-the-art methods on all three datasets.

READ FULL TEXT

page 1

page 8

research
10/26/2021

CTRN: Class-Temporal Relational Network for Action Detection

Action detection is an essential and challenging task, especially for de...
research
09/22/2022

FuTH-Net: Fusing Temporal Relations and Holistic Features for Aerial Video Classification

Unmanned aerial vehicles (UAVs) are now widely applied to data acquisiti...
research
10/17/2017

Single Shot Temporal Action Detection

Temporal action detection is a very important yet challenging problem, s...
research
12/31/2022

An end-to-end multi-scale network for action prediction in videos

In this paper, we develop an efficient multi-scale network to predict ac...
research
08/15/2023

Multi-scale Promoted Self-adjusting Correlation Learning for Facial Action Unit Detection

Facial Action Unit (AU) detection is a crucial task in affective computi...
research
08/07/2020

Multi-Level Temporal Pyramid Network for Action Detection

Currently, one-stage frameworks have been widely applied for temporal ac...
research
11/30/2021

Two-stage Temporal Modelling Framework for Video-based Depression Recognition using Graph Representation

Video-based automatic depression analysis provides a fast, objective and...

Please sign up or login with your details

Forgot password? Click here to reset