Video-Specific Query-Key Attention Modeling for Weakly-Supervised Temporal Action Localization

05/07/2023
by   Xijun Wang, et al.
0

Weakly-supervised temporal action localization aims to identify and localize the action instances in the untrimmed videos with only video-level action labels. When humans watch videos, we can adapt our abstract-level knowledge about actions in different video scenarios and detect whether some actions are occurring. In this paper, we mimic how humans do and bring a new perspective for locating and identifying multiple actions in a video. We propose a network named VQK-Net with a video-specific query-key attention modeling that learns a unique query for each action category of each input video. The learned queries not only contain the actions' knowledge features at the abstract level but also have the ability to fit this knowledge into the target video scenario, and they will be used to detect the presence of the corresponding action along the temporal dimension. To better learn these action category queries, we exploit not only the features of the current input video but also the correlation between different videos through a novel video-specific action category query learner worked with a query similarity loss. Finally, we conduct extensive experiments on three commonly used datasets (THUMOS14, ActivityNet1.2, and ActivityNet1.3) and achieve state-of-the-art performance.

READ FULL TEXT

page 1

page 8

research
01/21/2020

Weakly Supervised Temporal Action Localization Using Deep Metric Learning

Temporal action localization is an important step towards video understa...
research
04/25/2023

Weakly-Supervised Temporal Action Localization with Bidirectional Semantic Consistency Constraint

Weakly Supervised Temporal Action Localization (WTAL) aims to classify a...
research
04/28/2022

Tragedy Plus Time: Capturing Unintended Human Activities from Weakly-labeled Videos

In videos that contain actions performed unintentionally, agents do not ...
research
06/21/2022

Bi-Calibration Networks for Weakly-Supervised Video Representation Learning

The leverage of large volumes of web videos paired with the searched que...
research
04/06/2021

Zeus: Efficiently Localizing Actions in Videos using Reinforcement Learning

Detection and localization of actions in videos is an important problem ...
research
03/30/2023

JCDNet: Joint of Common and Definite phases Network for Weakly Supervised Temporal Action Localization

Weakly-supervised temporal action localization aims to localize action i...
research
11/19/2018

Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection

This paper proposes a segregated temporal assembly recurrent (STAR) netw...

Please sign up or login with your details

Forgot password? Click here to reset