Zero-Shot Audio Classification using Image Embeddings

06/10/2022
by   Duygu Dogan, et al.
0

Supervised learning methods can solve the given problem in the presence of a large set of labeled data. However, the acquisition of a dataset covering all the target classes typically requires manual labeling which is expensive and time-consuming. Zero-shot learning models are capable of classifying the unseen concepts by utilizing their semantic information. The present study introduces image embeddings as side information on zero-shot audio classification by using a nonlinear acoustic-semantic projection. We extract the semantic image representations from the Open Images dataset and evaluate the performance of the models on an audio subset of AudioSet using semantic information in different domains; image, audio, and textual. We demonstrate that the image embeddings can be used as semantic information to perform zero-shot audio classification. The experimental results show that the image and textual embeddings display similar performance both individually and together. We additionally calculate the semantic acoustic embeddings from the test samples to provide an upper limit to the performance. The results show that the classification performance is highly sensitive to the semantic relation between test and training classes and textual and image embeddings can reach up to the semantic acoustic embeddings when the seen and unseen classes are semantically similar.

READ FULL TEXT
research
05/06/2019

Zero-Shot Audio Classification Based on Class Label Embeddings

This paper proposes a zero-shot learning approach for audio classificati...
research
03/29/2021

A Simple Approach for Zero-Shot Learning based on Triplet Distribution Embeddings

Given the semantic descriptions of classes, Zero-Shot Learning (ZSL) aim...
research
09/15/2023

Exploring Meta Information for Audio-based Zero-shot Bird Classification

Advances in passive acoustic monitoring and machine learning have led to...
research
08/24/2022

Improved Zero-Shot Audio Tagging Classification with Patchout Spectrogram Transformers

Standard machine learning models for tagging and classifying acoustic si...
research
07/05/2019

Zero-shot Learning for Audio-based Music Classification and Tagging

Audio-based music classification and tagging is typically based on categ...
research
11/17/2018

Not just a matter of semantics: the relationship between visual similarity and semantic similarity

Knowledge transfer, zero-shot learning and semantic image retrieval are ...
research
07/15/2021

Semantic Image Cropping

Automatic image cropping techniques are commonly used to enhance the aes...

Please sign up or login with your details

Forgot password? Click here to reset