Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images

04/04/2015
by   Chen Sun, et al.
0

We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identify temporal segments corresponding to the actions, and learn models that generalize to unconstrained web videos. We find that web images queried by action names serve as well-localized highlights for many actions, but are noisily labeled. To solve this problem, we propose a simple yet effective method that takes weak video labels and noisy image labels as input, and generates localized action frames as output. This is achieved by cross-domain transfer between video frames and web images, using pre-trained deep convolutional neural networks. We then use the localized action frames to train action recognition models with long short-term memory networks. We collect a fine-grained sports action data set FGA-240 of more than 130,000 YouTube videos. It has 240 fine-grained actions under 85 sports activities. Convincing results are shown on the FGA-240 data set, as well as the THUMOS 2014 localization data set with untrimmed training videos.

READ FULL TEXT

page 1

page 3

page 4

page 7

page 8

research
04/19/2021

Temporal Query Networks for Fine-grained Video Understanding

Our objective in this work is fine-grained classification of actions in ...
research
12/22/2015

Do Less and Achieve More: Training CNNs for Action Recognition Utilizing Action Images from the Web

Recently, attempts have been made to collect millions of videos to train...
research
12/26/2017

SLAC: A Sparsely Labeled Dataset for Action Classification and Localization

This paper describes a procedure for the creation of large-scale video d...
research
06/09/2017

Learning to Learn from Noisy Web Videos

Understanding the simultaneously very diverse and intricately fine-grain...
research
03/28/2017

Towards Automatic Learning of Procedures from Web Instructional Videos

The potential for agents, whether embodied or software, to learn by obse...
research
08/18/2023

Progression-Guided Temporal Action Detection in Videos

We present a novel framework, Action Progression Network (APN), for temp...
research
07/21/2015

Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos

Every moment counts in action recognition. A comprehensive understanding...

Please sign up or login with your details

Forgot password? Click here to reset