Exploring CTC Based End-to-End Techniques for Myanmar Speech Recognition

05/13/2021

∙

In this work, we explore a Connectionist Temporal Classification (CTC) based end-to-end Automatic Speech Recognition (ASR) model for the Myanmar language. A series of experiments is presented on the topology of the model in which the convolutional layers are added and dropped, different depths of bidirectional long short-term memory (BLSTM) layers are used and different label encoding methods are investigated. The experiments are carried out in low-resource scenarios using our recorded Myanmar speech corpus of nearly 26 hours. The best model achieves character error rate (CER) of 4.72 (SER) of 12.38

READ FULL TEXT

Exploring CTC Based End-to-End Techniques for Myanmar Speech Recognition

Sign in with Google

Consider DeepAI Pro