Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events

04/10/2019
by   Hongwei Song, et al.
0

In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from existing strategies, which focus on characterizing global acoustical distributions of audio or the temporal evolution of short-term audio features, without analysis down to the level of sound events. To identify distinct sound events for each scene, we formulate ASC in a multi-instance learning (MIL) framework, where each audio recording is mapped into a bag-of-instances representation. Here, instances can be seen as high-level representations for sound events inside a scene. We also propose a MIL neural networks model, which implicitly identifies distinct instances (i.e., sound events). Furthermore, we propose two specially designed modules that model the multi-temporal scale and multi-modal natures of the sound events respectively. The experiments were conducted on the official development set of the DCASE2018 Task1 Subtask B, and our best-performing model improves over the official baseline by 9.4 This study indicates that recognizing acoustic scenes by identifying distinct sound events is effective and paves the way for future studies that combine this strategy with previous ones.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/14/2020

Sound Event Detection by Multitask Learning of Sound Events and Scenes with Soft Scene Labels

Sound event detection (SED) and acoustic scene classification (ASC) are ...
research
09/16/2019

Acoustic scene analysis with multi-head attention networks

Acoustic Scene Classification (ASC) is a challenging task, as a single s...
research
09/13/2022

Binaural Signal Representations for Joint Sound Event Detection and Acoustic Scene Classification

Sound event detection (SED) and Acoustic scene classification (ASC) are ...
research
04/26/2021

Identifying Actions for Sound Event Classification

In Psychology, actions are paramount for humans to perceive and separate...
research
03/16/2022

Instance-level loss based multiple-instance learning for acoustic scene classification

In acoustic scene classification (ASC) task, an acoustic scene consists ...
research
11/17/2018

The Intrinsic Memorability of Everyday Sounds

Our aural experience plays an integral role in the perception and memory...
research
12/11/2014

The bag-of-frames approach: a not so sufficient model for urban soundscapes

The "bag-of-frames" approach (BOF), which encodes audio signals as the l...

Please sign up or login with your details

Forgot password? Click here to reset