A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition

08/06/2020
by   Sitong Zhou, et al.
0

This paper presents a transfer learning method in speech emotion recognition based on a Time-Delay Neural Network (TDNN) architecture. A major challenge in the current speech-based emotion detection research is data scarcity. The proposed method resolves this problem by applying transfer learning techniques in order to leverage data from the automatic speech recognition (ASR) task for which ample data is available. Our experiments also show the advantage of speaker-class adaptation modeling techniques by adopting identity-vector (i-vector) based features in addition to standard Mel-Frequency Cepstral Coefficient (MFCC) features.[1] We show the transfer learning models significantly outperform the other methods without pretraining on ASR. The experiments performed on the publicly available IEMOCAP dataset which provides 12 hours of motional speech data. The transfer learning was initialized by using the Ted-Lium v.2 speech dataset providing 207 hours of audio with the corresponding transcripts. We achieve the highest significantly higher accuracy when compared to state-of-the-art, using five-fold cross validation. Using only speech, we obtain an accuracy 71.7 neutrality emotion content.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/23/2018

ASR-based Features for Emotion Recognition: A Transfer Learning Approach

During the last decade, the applications of signal processing have drast...
research
11/21/2019

Cantonese Automatic Speech Recognition Using Transfer Learning from Mandarin

We propose a system to develop a basic automatic speech recognizer(ASR) ...
research
10/26/2022

Pretrained audio neural networks for Speech emotion recognition in Portuguese

The goal of speech emotion recognition (SER) is to identify the emotiona...
research
03/09/2020

Deep Neural Networks for Automatic Speech Processing: A Survey from Large Corpora to Limited Data

Most state-of-the-art speech systems are using Deep Neural Networks (DNN...
research
03/27/2022

A Dataset for Speech Emotion Recognition in Greek Theatrical Plays

Machine learning methodologies can be adopted in cultural applications a...
research
07/04/2023

Pretraining Conformer with ASR or ASV for Anti-Spoofing Countermeasure

This paper introduces the Multi-scale Feature Aggregation Conformer (MFA...
research
07/21/2017

Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition

Unsupervised single-channel overlapped speech recognition is one of the ...

Please sign up or login with your details

Forgot password? Click here to reset