Deep Speaker Verification: Do We Need End to End?

06/22/2017
by   Dong Wang, et al.
0

End-to-end learning treats the entire system as a whole adaptable black box, which, if sufficient data are available, may learn a system that works very well for the target task. This principle has recently been applied to several prototype research on speaker verification (SV), where the feature learning and classifier are learned together with an objective function that is consistent with the evaluation metric. An opposite approach to end-to-end is feature learning, which firstly trains a feature learning model, and then constructs a back-end classifier separately to perform SV. Recently, both approaches achieved significant performance gains on SV, mainly attributed to the smart utilization of deep neural networks. However, the two approaches have not been carefully compared, and their respective advantages have not been well discussed. In this paper, we compare the end-to-end and feature learning approaches on a text-independent SV task. Our experiments on a dataset sampled from the Fisher database and involving 5,000 speakers demonstrated that the feature learning approach outperformed the end-to-end approach. This is a strong support for the feature learning approach, at least with data and computation resources similar to ours.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/10/2017

Deep Speaker Feature Learning for Text-independent Speaker Verification

Recently deep neural networks (DNNs) have been used to learn speaker fea...
research
07/31/2020

Feature Learning for Accelerometer based Gait Recognition

Recent advances in pattern matching, such as speech or object recognitio...
research
10/31/2017

Full-info Training for Deep Speaker Feature Learning

In recent studies, it has shown that speaker patterns can be learned fro...
research
06/22/2017

Cross-lingual Speaker Verification with Deep Feature Learning

Existing speaker verification (SV) systems often suffer from performance...
research
08/11/2017

Acoustic Feature Learning via Deep Variational Canonical Correlation Analysis

We study the problem of acoustic feature learning in the setting where w...
research
03/27/2022

End-to-End Active Speaker Detection

Recent advances in the Active Speaker Detection (ASD) problem build upon...
research
05/16/2022

Referring Expressions with Rational Speech Act Framework: A Probabilistic Approach

This paper focuses on a referring expression generation (REG) task in wh...

Please sign up or login with your details

Forgot password? Click here to reset