Bi-LSTM Scoring Based Similarity Measurement with Agglomerative Hierarchical Clustering (AHC) for Speaker Diarization

05/19/2022
by   Siddharth S. Nijhawan, et al.
0

Majority of speech signals across different scenarios are never available with well-defined audio segments containing only a single speaker. A typical conversation between two speakers consists of segments where their voices overlap, interrupt each other or halt their speech in between multiple sentences. Recent advancements in diarization technology leverage neural network-based approaches to improvise multiple subsystems of speaker diarization system comprising of extracting segment-wise embedding features and detecting changes in the speaker during conversation. However, to identify speaker through clustering, models depend on methodologies like PLDA to generate similarity measure between two extracted segments from a given conversational audio. Since these algorithms ignore the temporal structure of conversations, they tend to achieve a higher Diarization Error Rate (DER), thus leading to misdetections both in terms of speaker and change identification. Therefore, to compare similarity of two speech segments both independently and sequentially, we propose a Bi-directional Long Short-term Memory network for estimating the elements present in the similarity matrix. Once the similarity matrix is generated, Agglomerative Hierarchical Clustering (AHC) is applied to further identify speaker segments based on thresholding. To evaluate the performance, Diarization Error Rate (DER achieves a low DER of 34.80 Meeting Corpus as compared to traditional PLDA based similarity measurement mechanism which achieved a DER of 39.90

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/23/2019

LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization

More and more neural network approaches have achieved considerable impro...
research
10/22/2020

The HUAWEI Speaker Diarisation System for the VoxCeleb Speaker Diarisation Challenge

This paper describes system setup of our submission to speaker diarisati...
research
05/28/2023

Range-Based Equal Error Rate for Spoof Localization

Spoof localization, also called segment-level detection, is a crucial ta...
research
11/08/2022

BER: Balanced Error Rate For Speaker Diarization

DER is the primary metric to evaluate diarization performance while faci...
research
07/12/2019

Toeplitz Inverse Covariance based Robust Speaker Clustering for Naturalistic Audio Streams

Speaker diarization determines who spoke and when? in an audio stream. I...
research
04/06/2020

Probabilistic embeddings for speaker diarization

Speaker embeddings (x-vectors) extracted from very short segments of spe...
research
12/12/2017

Classification vs. Regression in Supervised Learning for Single Channel Speaker Count Estimation

The task of estimating the maximum number of concurrent speakers from si...

Please sign up or login with your details

Forgot password? Click here to reset