Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification

07/04/2019
by   Zhichao Zhang, et al.
0

Environmental sound classification (ESC) is a challenging problem due to the complexity of sounds. The ESC performance is heavily dependent on the effectiveness of representative features extracted from the environmental sounds. However, ESC often suffers from the semantically irrelevant frames and silent frames. In order to deal with this, we employ a frame-level attention model to focus on the semantically relevant frames and salient frames. Specifically, we first propose an convolutional recurrent neural network to learn spectro-temporal features and temporal correlations. Then, we extend our convolutional RNN model with a frame-level attention mechanism to learn discriminative feature representations for ESC. Experiments were conducted on ESC-50 and ESC-10 datasets. Experimental results demonstrated the effectiveness of the proposed method and achieved the state-of-the-art performance in terms of classification accuracy.

READ FULL TEXT

page 3

page 5

research
07/12/2020

Learning Frame Level Attention for Environmental Sound Classification

Environmental sound classification (ESC) is a challenging problem due to...
research
08/16/2019

Sub-Spectrogram Segmentation for Environmental Sound Classification via Convolutional Recurrent Neural Network and Score Level Fusion

Environmental Sound Classification (ESC) is an important and challenging...
research
05/28/2022

Feature Pyramid Attention based Residual Neural Network for Environmental Sound Classification

Environmental sound classification (ESC) is a challenging problem due to...
research
11/04/2020

A Multi-Channel Temporal Attention Convolutional Neural Network Model for Environmental Sound Classification

Recently, many attention-based deep neural networks have emerged and ach...
research
08/25/2018

Deep Convolutional Neural Network with Mixup for Environmental Sound Classification

Environmental sound classification (ESC) is an important and challenging...
research
03/12/2022

Recurrence-in-Recurrence Networks for Video Deblurring

State-of-the-art video deblurring methods often adopt recurrent neural n...
research
04/06/2019

Spatio-Temporal Attention Pooling for Audio Scene Classification

Acoustic scenes are rich and redundant in their content. In this work, w...

Please sign up or login with your details

Forgot password? Click here to reset