Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition

by   Antonio Miguel, et al.

In this paper we propose a method to model speaker and session variability and able to generate likelihood ratios using neural networks in an end-to-end phrase dependent speaker verification system. As in Joint Factor Analysis, the model uses tied hidden variables to model speaker and session variability and a MAP adaptation of some of the parameters of the model. In the training procedure our method jointly estimates the network parameters and the values of the speaker and channel hidden variables. This is done in a two-step backpropagation algorithm, first the network weights and factor loading matrices are updated and then the hidden variables, whose gradients are calculated by aggregating the corresponding speaker or session frames, since these hidden variables are tied. The last layer of the network is defined as a linear regression probabilistic model whose inputs are the previous layer outputs. This choice has the advantage that it produces likelihoods and additionally it can be adapted during the enrolment using MAP without the need of a gradient optimization. The decisions are made based on the ratio of the output likelihoods of two neural network models, speaker adapted and universal background model. The method was evaluated on the RSR2015 database.


page 1

page 2

page 3

page 4


Neural PLDA Modeling for End-to-End Speaker Verification

While deep learning models have made significant advances in supervised ...

Speaker Adaptive Training using Model Agnostic Meta-Learning

Speaker adaptive training (SAT) of neural network acoustic models learns...

Joint Speaker Encoder and Neural Back-end Model for Fully End-to-End Automatic Speaker Verification with Multiple Enrollment Utterances

Conventional automatic speaker verification systems can usually be decom...

Nondeterminism and Instability in Neural Network Optimization

Nondeterminism in neural network optimization produces uncertainty in pe...

PLDA with Two Sources of Inter-session Variability

In some speaker recognition scenarios we find conversations recorded sim...

Rapid Connectionist Speaker Adaptation

We present SVCnet, a system for modelling speaker variability. Encoder N...

Optimization of the Area Under the ROC Curve using Neural Network Supervectors for Text-Dependent Speaker Verification

This paper explores two techniques to improve the performance of text-de...

Please sign up or login with your details

Forgot password? Click here to reset