AFO-TAD: Anchor-free One-Stage Detector for Temporal Action Detection

10/18/2019
by   Yiping Tang, et al.
0

Temporal action detection is a fundamental yet challenging task in video understanding. Many of the state-of-the-art methods predict the boundaries of action instances based on predetermined anchors akin to the two-dimensional object detection detectors. However, it is hard to detect all the action instances with predetermined temporal scales because the durations of instances in untrimmed videos can vary from few seconds to several minutes. In this paper, we propose a novel action detection architecture named anchor-free one-stage temporal action detector (AFO-TAD). AFO-TAD achieves better performance for detecting action instances with arbitrary lengths and high temporal resolution, which can be attributed to two aspects. First, we design a receptive field adaption module which dynamically adjusts the receptive field for precise action detection. Second, AFO-TAD directly predicts the categories and boundaries at every temporal locations without predetermined anchors. Extensive experiments show that AFO-TAD improves the state-of-the-art performance on THUMOS'14.

READ FULL TEXT
research
06/29/2021

SRF-Net: Selective Receptive Field Network for Anchor-Free Temporal Action Detection

Temporal action detection (TAD) is a challenging task which aims to temp...
research
10/25/2022

Refining Action Boundaries for One-stage Detection

Current one-stage action detection methods, which simultaneously predict...
research
03/14/2022

RCL: Recurrent Continuous Localization for Temporal Action Detection

Temporal representation is the cornerstone of modern action detection te...
research
08/22/2020

Revisiting Anchor Mechanisms for Temporal Action Localization

Most of the current action localization methods follow an anchor-based p...
research
03/28/2023

STMixer: A One-Stage Sparse Action Detector

Traditional video action detectors typically adopt the two-stage pipelin...
research
07/05/2018

A Single Shot Text Detector with Scale-adaptive Anchors

Currently, most top-performing text detection networks tend to employ fi...
research
08/02/2019

Scale Matters: Temporal Scale Aggregation Network for Precise Action Localization in Untrimmed Videos

Temporal action localization is a recently-emerging task, aiming to loca...

Please sign up or login with your details

Forgot password? Click here to reset