Long-frame-shift Neural Speech Phase Prediction with Spectral Continuity Enhancement and Interpolation Error Compensation

08/17/2023
by   Yang Ai, et al.
0

Speech phase prediction, which is a significant research focus in the field of signal processing, aims to recover speech phase spectra from amplitude-related features. However, existing speech phase prediction methods are constrained to recovering phase spectra with short frame shifts, which are considerably smaller than the theoretical upper bound required for exact waveform reconstruction of short-time Fourier transform (STFT). To tackle this issue, we present a novel long-frame-shift neural speech phase prediction (LFS-NSPP) method which enables precise prediction of long-frame-shift phase spectra from long-frame-shift log amplitude spectra. The proposed method consists of three stages: interpolation, prediction and decimation. The short-frame-shift log amplitude spectra are first constructed from long-frame-shift ones through frequency-by-frequency interpolation to enhance the spectral continuity, and then employed to predict short-frame-shift phase spectra using an NSPP model, thereby compensating for interpolation errors. Ultimately, the long-frame-shift phase spectra are obtained from short-frame-shift ones through frame-by-frame decimation. Experimental results show that the proposed LFS-NSPP method can yield superior quality in predicting long-frame-shift phase spectra than the original NSPP model and other signal-processing-based phase estimation algorithms.

READ FULL TEXT

page 1

page 4

research
10/29/2018

STFT spectral loss for training a neural speech waveform model

This paper proposes a new loss using short-time Fourier transform (STFT)...
research
05/13/2023

APNet: An All-Frame-Level Neural Vocoder Incorporating Direct Prediction of Amplitude and Phase Spectra

This paper presents a novel neural vocoder named APNet which reconstruct...
research
01/05/2022

Frame Shift Prediction

Frame shift is a cross-linguistic phenomenon in translation which result...
research
03/02/2021

Signal recovery from a few linear measurements of its high-order spectra

The q-th order spectrum is a polynomial of degree q in the entries of a ...
research
11/29/2022

Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses

This paper presents a novel speech phase prediction model which predicts...
research
01/22/2016

A Robust Frame-based Nonlinear Prediction System for Automatic Speech Coding

In this paper, we propose a neural-based coding scheme in which an artif...
research
04/23/2018

The Future of Prosody: It's about Time

Prosody is usually defined in terms of the three distinct but interactin...

Please sign up or login with your details

Forgot password? Click here to reset