Meta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs

04/06/2020
by   Seong Min Kye, et al.
1

In realistic settings, a speaker recognition system needs to identify a speaker given a short utterance, while the utterance used to enroll may be relatively long. However, existing speaker recognition models perform poorly with such short utterances. To solve this problem, we introduce a meta-learning scheme with imbalance length pairs. Specifically, we use a prototypical network and train it with a support set of long utterances and a query set of short utterances. However, since optimizing for only the classes in the given episode is not sufficient to learn discriminative embeddings for other classes in the entire dataset, we additionally classify both support set and query set against the entire classes in the training set to learn a well-discriminated embedding space. By combining these two learning schemes, our model outperforms existing state-of-the-art speaker verification models learned in a standard supervised learning framework on short utterance (1-2 seconds) on VoxCeleb dataset. We also validate our proposed model for unseen speaker identification, on which it also achieves significant gain over existing approaches.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/31/2022

Improved Relation Networks for End-to-End Speaker Verification and Identification

Speaker identification systems in a real-world scenario are tasked to id...
research
05/07/2020

Crop Aggregating for short utterances speaker verification using raw waveforms

Most studies on speaker verification systems focus on long-duration utte...
research
10/16/2018

Deep neural network based i-vector mapping for speaker verification using short utterances

Text-independent speaker recognition using short utterances is a highly ...
research
04/01/2018

I-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker Verification

I-vector based text-independent speaker verification (SV) systems often ...
research
02/06/2019

Centroid-based deep metric learning for speaker recognition

Speaker embedding models that utilize neural networks to map utterances ...
research
03/29/2021

Improved Meta-learning training for Speaker Verification

Meta-learning (ML) has recently become a research hotspot in speaker ver...
research
02/26/2019

Utterance-level Aggregation For Speaker Recognition In The Wild

The objective of this paper is speaker recognition "in the wild"-where u...

Please sign up or login with your details

Forgot password? Click here to reset