Vietnamese Capitalization and Punctuation Recovery Models

07/04/2022
by   Hoang Thi Thu Uyen, et al.
0

Despite the rise of recent performant methods in Automatic Speech Recognition (ASR), such methods do not ensure proper casing and punctuation for their outputs. This problem has a significant impact on the comprehension of both Natural Language Processing (NLP) algorithms and human to process. Capitalization and punctuation restoration is imperative in pre-processing pipelines for raw textual inputs. For low resource languages like Vietnamese, public datasets for this task are scarce. In this paper, we contribute a public dataset for capitalization and punctuation recovery for Vietnamese; and propose a joint model for both tasks named JointCapPunc. Experimental results on the Vietnamese dataset show the effectiveness of our joint model compare to single model and previous joint learning model. We publicly release our dataset and the implementation of our model at https://github.com/anhtunguyen98/JointCapPunc

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/30/2021

How Low is Too Low? A Computational Perspective on Extremely Low-Resource Languages

Despite the recent advancements of attention-based deep learning archite...
research
11/21/2021

Capitalization and Punctuation Restoration: a Survey

Ensuring proper punctuation and letter casing is a key pre-processing st...
research
10/25/2019

L2RS: A Learning-to-Rescore Mechanism for Automatic Speech Recognition

Modern Automatic Speech Recognition (ASR) systems primarily rely on scor...
research
12/22/2020

Adversarial Meta Sampling for Multilingual Low-Resource Speech Recognition

Low-resource automatic speech recognition (ASR) is challenging, as the l...
research
11/06/2021

Towards Building ASR Systems for the Next Billion Users

Recent methods in speech and language technology pretrain very LARGE mod...
research
09/06/2023

RoDia: A New Dataset for Romanian Dialect Identification from Speech

Dialect identification is a critical task in speech processing and langu...

Please sign up or login with your details

Forgot password? Click here to reset