Learning Sparse Analytic Filters for Piano Transcription

08/23/2021
by   Frank Cwitkowitz, et al.
7

In recent years, filterbank learning has become an increasingly popular strategy for various audio-related machine learning tasks. This is partly due to its ability to discover task-specific audio characteristics which can be leveraged in downstream processing. It is also a natural extension of the nearly ubiquitous deep learning methods employed to tackle a diverse array of audio applications. In this work, several variations of a frontend filterbank learning module are investigated for piano transcription, a challenging low-level music information retrieval task. We build upon a standard piano transcription model, modifying only the feature extraction stage. The filterbank module is designed such that its complex filters are unconstrained 1D convolutional kernels with long receptive fields. Additional variations employ the Hilbert transform to render the filters intrinsically analytic and apply variational dropout to promote filterbank sparsity. Transcription results are compared across all experiments, and we offer visualization and analysis of the filterbanks.

READ FULL TEXT

page 11

page 13

page 15

page 17

page 19

page 21

page 23

page 25

research
06/15/2019

Audio-Based Music Classification with DenseNet And Data Augmentation

In recent years, deep learning technique has received intense attention ...
research
12/10/2022

A Comparison of Audio Preprocessing Techniques and Deep Learning Algorithms for Raga Recognition

Ragas form the foundation for Indian Classical Music. The task of Raga R...
research
01/21/2021

Effect of Deep Learning Feature Inference Techniques on Respiratory Sounds

Analysis of respiratory sounds increases its importance every day. Many ...
research
02/26/2023

From Audio to Symbolic Encoding

Automatic music transcription (AMT) aims to convert raw audio to symboli...
research
11/27/2018

Learning to detect dysarthria from raw speech

Speech classifiers of paralinguistic traits traditionally learn from div...
research
03/30/2023

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

The advancement of audio-language (AL) multimodal learning tasks has bee...

Please sign up or login with your details

Forgot password? Click here to reset