Audio-noise Power Spectral Density Estimation Using Long Short-term Memory

04/10/2019
by   Xiaofei Li, et al.
0

We propose a method using a long short-term memory (LSTM) network to estimate the noise power spectral density (PSD) of single-channel audio signals represented in the short time Fourier transform (STFT) domain. An LSTM network common to all frequency bands is trained, which processes each frequency band individually by mapping the noisy STFT magnitude sequence to its corresponding noise PSD sequence. Unlike deep-learning-based speech enhancement methods that learn the full-band spectral structure of speech segments, the proposed method exploits the sub-band STFT magnitude evolution of noise with a long time dependency, in the spirit of the unsupervised noise estimators described in the literature. Speaker- and speech-independent experiments with different types of noise show that the proposed method outperforms the unsupervised estimators, and generalizes well to noise types that are not present in the training set.

READ FULL TEXT
research
11/25/2019

Narrow-band Deep Filtering for Multichannel Speech Enhancement

In this paper we address the problem of multichannel speech enhancement ...
research
11/11/2019

Supervised Initialization of LSTM Networks for Fundamental Frequency Detection in Noisy Speech Signals

Fundamental frequency is one of the most important parameters of human s...
research
03/23/2022

FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement

Previously proposed FullSubNet has achieved outstanding performance in D...
research
02/23/2023

Frequency bin-wise single channel speech presence probability estimation using multiple DNNs

In this work, we propose a frequency bin-wise method to estimate the sin...
research
05/24/2018

VisemeNet: Audio-Driven Animator-Centric Speech Animation

We present a novel deep-learning based approach to producing animator-ce...
research
04/18/2023

Neural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated Full- and Sub-Band Modeling

We propose FSB-LSTM, a novel long short-term memory (LSTM) based archite...
research
12/04/2018

LSTM based AE-DNN constraint for better late reverb suppression in multi-channel LP formulation

Prediction of late reverberation component using multi-channel linear pr...

Please sign up or login with your details

Forgot password? Click here to reset