Meta-learning with Latent Space Clustering in Generative Adversarial Network for Speaker Diarization

07/19/2020
by   Monisankha Pal, et al.
0

The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an encoder network (ClusterGAN) to project input x-vectors into a latent space has shown promising performance on meeting data. In this paper, we extend the ClusterGAN network to improve diarization robustness and enable rapid generalization across various challenging domains. To this end, we fetch the pre-trained encoder from the ClusterGAN and fine-tune it by using prototypical loss (meta-ClusterGAN or MCGAN) under the meta-learning paradigm. Experiments are conducted on CALLHOME telephonic conversations, AMI meeting data, DIHARD II (dev set) which includes challenging multi-domain corpus, and two child-clinician interaction corpora (ADOS, BOSCC) related to the autism spectrum disorder domain. Extensive analyses of the experimental data are done to investigate the effectiveness of the proposed ClusterGAN and MCGAN embeddings over x-vectors. The results show that the proposed embeddings with normalized maximum eigengap spectral clustering (NME-SC) back-end consistently outperform Kaldi state-of-the-art z-vector diarization system. Finally, we employ embedding fusion with x-vectors to provide further improvement in diarization performance. We achieve a relative diarization error rate (DER) improvement of 6.67 proposed fused embeddings over x-vectors. Besides, the MCGAN embeddings provide better performance in the number of speakers estimation and short speech segment diarization as compared to x-vectors and ClusterGAN in telephonic data.

READ FULL TEXT

page 1

page 9

page 10

research
10/24/2019

Speaker diarization using latent space clustering in generative adversarial network

In this work, we propose deep latent space clustering for speaker diariz...
research
07/31/2020

Designing Neural Speaker Embeddings with Meta Learning

Neural speaker embeddings trained using classification objectives have d...
research
10/24/2019

Meta-learning for robust child-adult classification from speech

Computational modeling of naturalistic conversations in clinical applica...
research
11/07/2018

Generative Adversarial Speaker Embedding Networks for Domain Robust End-to-End Speaker Verification

This article presents a novel approach for learning domain-invariant spe...
research
10/24/2019

A study of semi-supervised speaker diarization system using gan mixture model

We propose a new speaker diarization system based on a recently introduc...
research
10/22/2020

Combination of Deep Speaker Embeddings for Diarisation

Recently, significant progress has been made in speaker diarisation afte...
research
06/22/2023

Implicit spoken language diarization

Spoken language diarization (LD) and related tasks are mostly explored u...

Please sign up or login with your details

Forgot password? Click here to reset