Deep-Learning-Based Audio-Visual Speech Enhancement in Presence of Lombard Effect

05/29/2019
by   Daniel Michelsanti, et al.
0

When speaking in presence of background noise, humans reflexively change their way of speaking in order to improve the intelligibility of their speech. This reflex is known as Lombard effect. Collecting speech in Lombard conditions is usually hard and costly. For this reason, speech enhancement systems are generally trained and evaluated on speech recorded in quiet to which noise is artificially added. Since these systems are often used in situations where Lombard speech occurs, in this work we perform an analysis of the impact that Lombard effect has on audio, visual and audio-visual speech enhancement, focusing on deep-learning-based systems, since they represent the current state of the art in the field. We conduct several experiments using an audio-visual Lombard speech corpus consisting of utterances spoken by 54 different talkers. The results show that training deep-learning-based models with Lombard speech is beneficial in terms of both estimated speech quality and estimated speech intelligibility at low signal to noise ratios, where the visual modality can play an important role in acoustically challenging situations. We also find that a performance difference between genders exists due to the distinct Lombard speech exhibited by males and females, and we analyse it in relation with acoustic and visual features. Furthermore, listening tests conducted with audio-visual stimuli show that the speech quality of the signals processed with systems trained using Lombard speech is statistically significantly better than the one obtained using systems trained with non-Lombard speech at a signal to noise ratio of -5 dB. Regarding speech intelligibility, we find a general tendency of the benefit in training the systems with Lombard speech.

READ FULL TEXT

page 6

page 9

research
11/15/2018

Effects of Lombard Reflex on the Performance of Deep-Learning-Based Audio-Visual Speech Enhancement Systems

Humans tend to change their way of speaking when they are immersed in a ...
research
08/21/2020

An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

Speech enhancement and speech separation are two related tasks, whose pu...
research
07/04/2017

Hidden-Markov-Model Based Speech Enhancement

The goal of this contribution is to use a parametric speech synthesis sy...
research
11/15/2018

On Training Targets and Objective Functions for Deep-Learning-Based Audio-Visual Speech Enhancement

Audio-visual speech enhancement (AV-SE) is the task of improving speech ...
research
11/04/2022

Speech enhancement using ego-noise references with a microphone array embedded in an unmanned aerial vehicle

A method is proposed for performing speech enhancement using ego-noise r...
research
06/04/2021

A Database for Research on Detection and Enhancement of Speech Transmitted over HF links

In this paper we present an open database for the development of detecti...
research
02/10/2020

On Cross-Corpus Generalization of Deep Learning Based Speech Enhancement

In recent years, supervised approaches using deep neural networks (DNNs)...

Please sign up or login with your details

Forgot password? Click here to reset