Log In Sign Up

Independent Deeply Learned Tensor Analysis for Determined Audio Source Separation

by   Naoki Narisawa, et al.

We address the determined audio source separation problem in the time-frequency domain. In independent deeply learned matrix analysis (IDLMA), it is assumed that the inter-frequency correlation of each source spectrum is zero, which is inappropriate for modeling nonstationary signals such as music signals. To account for the correlation between frequencies, independent positive semidefinite tensor analysis has been proposed. This unsupervised (blind) method, however, severely restrict the structure of frequency covariance matrices (FCMs) to reduce the number of model parameters. As an extension of these conventional approaches, we here propose a supervised method that models FCMs using deep neural networks (DNNs). It is difficult to directly infer FCMs using DNNs. Therefore, we also propose a new FCM model represented as a convex combination of a diagonal FCM and a rank-1 FCM. Our FCM model is flexible enough to not only consider inter-frequency correlation, but also capture the dynamics of time-varying FCMs of nonstationary signals. We infer the proposed FCMs using two DNNs: DNN for power spectrum estimation and DNN for time-domain signal estimation. An experimental result of separating music signals shows that the proposed method provides higher separation performance than IDLMA.


Independent Deeply Learned Matrix Analysis for Multichannel Audio Source Separation

In this paper, we address a multichannel audio source separation task an...

Sampling-Frequency-Independent Audio Source Separation Using Convolution Layer Based on Impulse Invariant Method

Audio source separation is often used as preprocessing of various applic...

All for One and One for All: Improving Music Separation by Bridging Networks

This paper proposes several improvements for music separation with deep ...

Sampling Frequency Independent Dialogue Separation

In some DNNs for audio source separation, the relevant model parameters ...

Empirical Bayesian Independent Deeply Learned Matrix Analysis For Multichannel Audio Source Separation

Independent deeply learned matrix analysis (IDLMA) is one of the state-o...

Determined BSS based on time-frequency masking and its application to harmonic vector analysis

When the number of microphones is equal to that of the source signals (t...

A context encoder for audio inpainting

We studied the ability of deep neural networks (DNNs) to restore missing...