Multi-Language Identification Using Convolutional Recurrent Neural Network

11/12/2016
by   Vrishabh Ajay Lakhani, et al.
0

Language Identification, being an important aspect of Automatic Speaker Recognition has had many changes and new approaches to ameliorate performance over the last decade. We compare the performance of using audio spectrum in the log scale and using Polyphonic sound sequences from raw audio samples to train the neural network and to classify speech as either English or Spanish. To achieve this, we use the novel approach of using a Convolutional Recurrent Neural Network using Long Short Term Memory (LSTM) or a Gated Recurrent Unit (GRU) for forward propagation of the neural network. Our hypothesis is that the performance of using polyphonic sound sequence as features and both LSTM and GRU as the gating mechanisms for the neural network outperform the traditional MFCC features using a unidirectional Deep Neural Network.

READ FULL TEXT
research
12/11/2014

Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

In this paper we compare different types of recurrent units in recurrent...
research
12/19/2019

LSTM-TDNN with convolutional front-end for Dialect Identification in the 2019 Multi-Genre Broadcast Challenge

This paper presents a novel Dialect Identification (DID) system develope...
research
07/25/2017

SAR Target Recognition Using the Multi-aspect-aware Bidirectional LSTM Recurrent Neural Networks

The outstanding pattern recognition performance of deep learning brings ...
research
07/03/2018

Weakly Supervised Deep Recurrent Neural Networks for Basic Dance Step Generation

A deep recurrent neural network with audio input is applied to model bas...
research
01/27/2022

Deep Recurrent Learning for Heart Sounds Segmentation based on Instantaneous Frequency Features

In this work, a novel stack of well-known technologies is presented to d...
research
06/04/2021

Classification of Audio Segments in Call Center Recordings using Convolutional Recurrent Neural Networks

Detailed statistical analysis of call center recordings is critical in t...
research
10/29/2021

Personalized breath based biometric authentication with wearable multimodality

Breath with nose sound features has been shown as a potential biometric ...

Please sign up or login with your details

Forgot password? Click here to reset