Generalizing AUC Optimization to Multiclass Classification for Audio Segmentation With Limited Training Data

10/27/2021
by   Pablo Gimeno, et al.
16

Area under the ROC curve (AUC) optimisation techniques developed for neural networks have recently demonstrated their capabilities in different audio and speech related tasks. However, due to its intrinsic nature, AUC optimisation has focused only on binary tasks so far. In this paper, we introduce an extension to the AUC optimisation framework so that it can be easily applied to an arbitrary number of classes, aiming to overcome the issues derived from training data limitations in deep learning solutions. Building upon the multiclass definitions of the AUC metric found in the literature, we define two new training objectives using a one-versus-one and a one-versus-rest approach. In order to demonstrate its potential, we apply them in an audio segmentation task with limited training data that aims to differentiate 3 classes: foreground music, background music and no music. Experimental results show that our proposal can improve the performance of audio segmentation systems significantly compared to traditional training criteria such as cross entropy.

READ FULL TEXT
research
08/30/2021

Unsupervised Learning of Deep Features for Music Segmentation

Music segmentation refers to the dual problem of identifying boundaries ...
research
06/08/2020

A Modified AUC for Training Convolutional Neural Networks: Taking Confidence into Account

Receiver operating characteristic (ROC) curve is an informative tool in ...
research
07/13/2021

AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data

Deep neural networks provide effective solutions to small-footprint keyw...
research
02/09/2021

Enhancing Audio Augmentation Methods with Consistency Learning

Data augmentation is an inexpensive way to increase training data divers...
research
08/03/2023

MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Diffusion models have shown promising results in cross-modal generation ...
research
11/15/2022

SSM-Net: feature learning for Music Structure Analysis using a Self-Similarity-Matrix based loss

In this paper, we propose a new paradigm to learn audio features for Mus...

Please sign up or login with your details

Forgot password? Click here to reset