Efficient Learning of Harmonic Priors for Pitch Detection in Polyphonic Music

05/19/2017
by   Pablo A. Alvarado, et al.
0

Automatic music transcription (AMT) aims to infer a latent symbolic representation of a piece of music (piano-roll), given a corresponding observed audio recording. Transcribing polyphonic music (when multiple notes are played simultaneously) is a challenging problem, due to highly structured overlapping between harmonics. We study whether the introduction of physically inspired Gaussian process (GP) priors into audio content analysis models improves the extraction of patterns required for AMT. Audio signals are described as a linear combination of sources. Each source is decomposed into the product of an amplitude-envelope, and a quasi-periodic component process. We introduce the Matérn spectral mixture (MSM) kernel for describing frequency content of singles notes. We consider two different regression approaches. In the sigmoid model every pitch-activation is independently non-linear transformed. In the softmax model several activation GPs are jointly non-linearly transformed. This introduce cross-correlation between activations. We use variational Bayes for approximate inference. We empirically evaluate how these models work in practice transcribing polyphonic music. We demonstrate that rather than encourage dependency between activations, what is relevant for improving pitch detection is to learnt priors that fit the frequency content of the sound events to detect.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/03/2016

Gaussian Processes for Music Audio Modelling and Content Analysis

Real music signals are highly variable, yet they have strong statistical...
research
01/31/2019

End-to-End Probabilistic Inference for Nonstationary Audio Analysis

A typical audio signal processing pipeline includes multiple disjoint an...
research
12/15/2016

Towards Score Following in Sheet Music Images

This paper addresses the matching of short music audio snippets to the c...
research
10/14/2021

Student-t Networks for Melody Estimation

Melody estimation or melody extraction refers to the extraction of the p...
research
10/30/2018

Sparse Gaussian process Audio Source Separation Using Spectrum Priors in the Time-Domain

Gaussian process (GP) audio source separation is a time-domain approach ...
research
03/28/2017

Particle Filtering for PLCA model with Application to Music Transcription

Automatic Music Transcription (AMT) consists in automatically estimating...
research
06/28/2022

Volume-Independent Music Matching by Frequency Spectrum Comparison

Often, I hear a piece of music and wonder what the name of the piece is....

Please sign up or login with your details

Forgot password? Click here to reset