Self-supervised Representation Learning With Path Integral Clustering For Speaker Diarization

04/19/2021
by   Prachi Singh, et al.
0

Automatic speaker diarization techniques typically involve a two-stage processing approach where audio segments of fixed duration are converted to vector representations in the first stage. This is followed by an unsupervised clustering of the representations in the second stage. In most of the prior approaches, these two stages are performed in an isolated manner with independent optimization steps. In this paper, we propose a representation learning and clustering algorithm that can be iteratively performed for improved speaker diarization. The representation learning is based on principles of self-supervised learning while the clustering algorithm is a graph structural method based on path integral clustering (PIC). The representation learning step uses the cluster targets from PIC and the clustering step is performed on embeddings learned from the self-supervised deep model. This iterative approach is referred to as self-supervised clustering (SSC). The diarization experiments are performed on CALLHOME and AMI meeting datasets. In these experiments, we show that the SSC algorithm improves significantly over the baseline system (relative improvements of 13 CALLHOME and AMI datasets respectively in terms of diarization error rate (DER)). In addition, the DER results reported in this work improve over several other recent approaches for speaker diarization.

READ FULL TEXT

page 1

page 2

page 4

page 7

page 8

page 11

research
09/14/2021

Self-Supervised Metric Learning With Graph Clustering For Speaker Diarization

In this paper, we propose a novel algorithm for speaker diarization usin...
research
10/25/2020

An iterative framework for self-supervised deep speaker representation learning

In this paper, we propose an iterative framework for self-supervised spe...
research
08/09/2023

Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation

The emergence of self-supervised representation (i.e., wav2vec 2.0) allo...
research
02/24/2023

Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization

Conventional methods for speaker diarization involve windowing an audio ...
research
02/27/2020

GATCluster: Self-Supervised Gaussian-Attention Network for Image Clustering

Deep clustering has achieved state-of-the-art results via joint represen...
research
05/28/2018

Resolving Event Coreference with Supervised Representation Learning and Clustering-Oriented Regularization

We present an approach to event coreference resolution by developing a g...
research
02/14/2022

Learning Weakly-Supervised Contrastive Representations

We argue that a form of the valuable information provided by the auxilia...

Please sign up or login with your details

Forgot password? Click here to reset