Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

10/28/2020
by   Tsun-An Hsieh, et al.
0

Speech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g. phones and syllables. In this study, we took phonetic characteristics into account in the SE training process. Hence, we designed a phone-fortified perceptual (PFP) loss, and the training of our SE model was guided by PFP loss. In PFP loss, phonetic characteristics are extracted by wav2vec, an unsupervised learning model based on the contrastive predictive coding (CPC) criterion. Different from previous deep-feature-based approaches, the proposed approach explicitly uses the phonetic information in the deep feature extraction process to guide the SE model training. To test the proposed approach, we first confirmed that the wav2vec representations carried clear phonetic information using a t-distributed stochastic neighbor embedding (t-SNE) analysis. Next, we observed that the proposed PFP loss was more strongly correlated with the perceptual evaluation metrics than point-wise and signal-level losses, thus achieving higher scores for standardized quality and intelligibility evaluation metrics in the Voice Bank-DEMAND dataset.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/31/2022

Perceptual Contrast Stretching on Target Feature for Speech Enhancement

Speech enhancement (SE) performance has improved considerably since the ...
research
11/15/2020

Speech enhancement guided by contextual articulatory information

Previous studies have confirmed the effectiveness of leveraging articula...
research
10/19/2021

Speech Enhancement Based on Cyclegan with Noise-informed Training

Speech enhancement (SE) approaches can be classified into supervised and...
research
11/03/2021

Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features

In this study, we propose a cross-domain multi-objective speech assessme...
research
10/20/2020

Investigating Cross-Domain Losses for Speech Enhancement

Recent years have seen a surge in the number of available frameworks for...
research
08/13/2020

Incorporating Broad Phonetic Information for Speech Enhancement

In noisy conditions, knowing speech contents facilitates listeners to mo...
research
11/16/2022

A Two-Stage Deep Representation Learning-Based Speech Enhancement Method Using Variational Autoencoder and Adversarial Training

This paper focuses on leveraging deep representation learning (DRL) for ...

Please sign up or login with your details

Forgot password? Click here to reset