Temporal Bilinear Encoding Network of Audio-Visual Features at Low Sampling Rates

12/18/2020
by   Feiyan Hu, et al.
1

Current deep learning based video classification architectures are typically trained end-to-end on large volumes of data and require extensive computational resources. This paper aims to exploit audio-visual information in video classification with a 1 frame per second sampling rate. We propose Temporal Bilinear Encoding Networks (TBEN) for encoding both audio and visual long range temporal information using bilinear pooling and demonstrate bilinear pooling is better than average pooling on the temporal dimension for videos with low sampling rate. We also embed the label hierarchy in TBEN to further improve the robustness of the classifier. Experiments on the FGA240 fine-grained classification dataset using TBEN achieve a new state-of-the-art (hit@1=47.95 multiple decoupled modalities like visual semantic and motion features: experiments on UCF101 sampled at 1 FPS achieve close to state-of-the-art accuracy (hit@1=91.03 resources than competing approaches for both training and prediction.

READ FULL TEXT
POST COMMENT

Comments

There are no comments yet.

Authors

page 1

page 2

page 3

page 4

07/26/2018

Hierarchical Bilinear Pooling for Fine-Grained Visual Recognition

Fine-grained visual recognition is challenging because it highly relies ...
09/16/2018

Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

Leveraging both visual frames and audio has been experimentally proven e...
07/25/2020

Approximated Bilinear Modules for Temporal Modeling

We consider two less-emphasized temporal properties of video: 1. Tempora...
11/16/2016

Low-rank Bilinear Pooling for Fine-Grained Classification

Pooling second-order local feature statistics to form a high-dimensional...
12/05/2018

Local Temporal Bilinear Pooling for Fine-grained Action Parsing

Fine-grained temporal action parsing is important in many applications, ...
05/23/2017

Look, Listen and Learn

We consider the question: what can be learnt by looking at and listening...
08/26/2019

CoinNet: Deep Ancient Roman Republican Coin Classification via Feature Fusion and Attention

We perform classification of ancient Roman Republican coins via recogniz...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.