Dual Normalization Multitasking for Audio-Visual Sounding Object Localization

06/01/2021
by   Tokuhiro Nishikawa, et al.
0

Although several research works have been reported on audio-visual sound source localization in unconstrained videos, no datasets and metrics have been proposed in the literature to quantitatively evaluate its performance. Defining the ground truth for sound source localization is difficult, because the location where the sound is produced is not limited to the range of the source object, but the vibrations propagate and spread through the surrounding objects. Therefore we propose a new concept, Sounding Object, to reduce the ambiguity of the visual location of sound, making it possible to annotate the location of the wide range of sound sources. With newly proposed metrics for quantitative evaluation, we formulate the problem of Audio-Visual Sounding Object Localization (AVSOL). We also created the evaluation dataset (AVSOL-E dataset) by manually annotating the test set of well-known Audio-Visual Event (AVE) dataset. To tackle this new AVSOL problem, we propose a novel multitask training strategy and architecture called Dual Normalization Multitasking (DNM), which aggregates the Audio-Visual Correspondence (AVC) task and the classification task for video events into a single audio-visual similarity map. By efficiently utilize both supervisions by DNM, our proposed architecture significantly outperforms the baseline methods.

READ FULL TEXT

page 1

page 7

page 8

research
08/30/2022

A Closer Look at Weakly-Supervised Audio-Visual Source Localization

Audio-visual source localization is a challenging task that aims to pred...
research
03/17/2022

Localizing Visual Sounds the Easy Way

Unsupervised audio-visual source localization aims at localizing visible...
research
06/15/2023

STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

While direction of arrival (DOA) of sound events is generally estimated ...
research
03/23/2023

Egocentric Audio-Visual Object Localization

Humans naturally perceive surrounding scenes by unifying sound and sight...
research
04/11/2022

How to Listen? Rethinking Visual Sound Localization

Localizing visual sounds consists on locating the position of objects th...
research
04/11/2019

The Sound of Motions

Sounds originate from object motions and vibrations of surrounding air. ...
research
08/08/2023

Dual input neural networks for positional sound source localization

In many signal processing applications, metadata may be advantageously u...

Please sign up or login with your details

Forgot password? Click here to reset