Uniphore's submission to Fearless Steps Challenge Phase-2

06/10/2020
by   Karthik Pandia D S, et al.
0

We propose supervised systems for speech activity detection (SAD) and speaker identification (SID) tasks in Fearless Steps Challenge Phase-2. The proposed systems for both the tasks share a common convolutional neural network (CNN) architecture. Mel spectrogram is used as features. For speech activity detection, the spectrogram is divided into smaller overlapping chunks. The network is trained to recognize the chunks. The network architecture and the training steps used for the SID task are similar to that of the SAD task, except that longer spectrogram chunks are used. We propose a two-level identification method for SID task. First, for each chunk, a set of speakers is hypothesized based on the neural network posterior probabilities. Finally, the speaker identity of the utterance is identified using the chunk-level hypotheses by applying a voting rule. On SAD task, a detection cost function score of 5.96 top 5 retrieval accuracy of 82.07 sets for SID task. A brief analysis is made on the results to provide insights into the miss-classified cases in both the tasks.

READ FULL TEXT

page 2

page 3

research
06/23/2021

Enrollment-less training for personalized voice activity detection

We present a novel personalized voice activity detection (PVAD) learning...
research
09/20/2022

The BUCEA Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2022

This paper describes the BUCEA speaker diarization system for the 2022 V...
research
11/15/2022

Is Style All You Need? Dependencies Between Emotion and GST-based Speaker Recognition

In this work, we study the hypothesis that speaker identity embeddings e...
research
10/30/2021

Real-time Speaker counting in a cocktail party scenario using Attention-guided Convolutional Neural Network

Most current speech technology systems are designed to operate well even...
research
04/07/2021

Three-class Overlapped Speech Detection using a Convolutional Recurrent Neural Network

In this work, we propose an overlapped speech detection system trained a...
research
11/03/2022

Fearless Steps Challenge Phase-1 Evaluation Plan

The Fearless Steps Challenge 2019 Phase-1 (FSC-P1) is the inaugural Chal...
research
01/24/2020

Character-independent font identification

There are a countless number of fonts with various shapes and styles. In...

Please sign up or login with your details

Forgot password? Click here to reset