Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization

02/24/2023
by   Prachi Singh, et al.
0

Conventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multi-step approach generates speaker assignments for each segment. In this paper, we propose a novel Supervised HierArchical gRaph Clustering algorithm (SHARC) for speaker diarization where we introduce a hierarchical structure using Graph Neural Network (GNN) to perform supervised clustering. The supervision allows the model to update the representations and directly improve the clustering performance, thus enabling a single-step approach for diarization. In the proposed work, the input segment embeddings are treated as nodes of a graph with the edge weights corresponding to the similarity scores between the nodes. We also propose an approach to jointly update the embedding extractor and the GNN model to perform end-to-end speaker diarization (E2E-SHARC). During inference, the hierarchical clustering is performed using node densities and edge existence probabilities to merge the segments until convergence. In the diarization experiments, we illustrate that the proposed E2E-SHARC approach achieves 53 the baseline systems on benchmark datasets like AMI and Voxconverse, respectively.

READ FULL TEXT
research
09/14/2021

Self-Supervised Metric Learning With Graph Clustering For Speaker Diarization

In this paper, we propose a novel algorithm for speaker diarization usin...
research
07/03/2021

Learning Hierarchical Graph Neural Networks for Image Clustering

We propose a hierarchical graph neural network (GNN) model that learns h...
research
05/22/2020

Speaker diarization with session-level speaker embedding refinement using graph neural networks

Deep speaker embedding models have been commonly used as a building bloc...
research
04/19/2021

Self-supervised Representation Learning With Path Integral Clustering For Speaker Diarization

Automatic speaker diarization techniques typically involve a two-stage p...
research
05/23/2023

Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization

Combining end-to-end neural speaker diarization (EEND) with vector clust...
research
10/13/2021

SSSNET: Semi-Supervised Signed Network Clustering

Node embeddings are a powerful tool in the analysis of networks; yet, th...
research
04/26/2022

Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure

Clustering-based speaker diarization has stood firm as one of the major ...

Please sign up or login with your details

Forgot password? Click here to reset