Multi-input Architecture and Disentangled Representation Learning for Multi-dimensional Modeling of Music Similarity

11/02/2021
by   Sebastian Ribecky, et al.
0

In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example scenario. Music however, naturally decomposes into a set of semantically meaningful factors of variation. Current representation learning strategies pursue the disentanglement of such factors from deep representations, resulting in highly interpretable models. This allows the modeling of music similarity perception, which is highly subjective and multi-dimensional. While the focus of prior work is on metadata driven notions of similarity, we suggest to directly model the human notion of multi-dimensional music similarity. To achieve this, we propose a multi-input deep neural network architecture, which simultaneously processes mel-spectrogram, CENS-chromagram and tempogram in order to extract informative features for the different disentangled musical dimensions: genre, mood, instrument, era, tempo, and key. We evaluated the proposed music similarity approach using a triplet prediction task and found that the proposed multi-input architecture outperforms a state of the art method. Furthermore, we present a novel multi-dimensional analysis in order to evaluate the influence of each disentangled dimension on the perception of music similarity.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/09/2020

Metric Learning vs Classification for Disentangled Music Representation Learning

Deep representation learning offers a powerful paradigm for mapping inpu...
research
08/09/2020

Disentangled Multidimensional Metric Learning for Music Similarity

Music similarity search is useful for a variety of creative tasks such a...
research
11/08/2018

Learning Disentangled Representations for Timber and Pitch in Music Audio

Timbre and pitch are the two main perceptual properties of musical sound...
research
07/19/2023

DisCover: Disentangled Music Representation Learning for Cover Song Identification

In the field of music information retrieval (MIR), cover song identifica...
research
09/17/2019

Multi-Task Music Representation Learning from Multi-Label Embeddings

This paper presents a novel approach to music representation learning. T...
research
06/13/2022

Self-Supervised Representation Learning With MUlti-Segmental Informational Coding (MUSIC)

Self-supervised representation learning maps high-dimensional data into ...
research
03/23/2021

A General Framework for Learning Prosodic-Enhanced Representation of Rap Lyrics

Learning and analyzing rap lyrics is a significant basis for many web ap...

Please sign up or login with your details

Forgot password? Click here to reset