A Pilot Study on Multiple Choice Machine Reading Comprehension for Vietnamese Texts

01/16/2020
by   Kiet Van Nguyen, et al.
0

Machine Reading Comprehension (MRC) is the task of natural language processing which studies the ability to read and understand unstructured texts and then find the correct answers for questions. Until now, we have not yet had any MRC dataset for such a low-resource language as Vietnamese. In this paper, we introduce ViMMRC, a challenging machine comprehension corpus with multiple-choice questions, intended for research on the machine comprehension of Vietnamese text. This corpus includes 2,783 multiple-choice questions and answers based on a set of 417 Vietnamese texts used for teaching reading comprehension for 1st to 5th graders. Answers may be extracted from the contents of single or multiple sentences in the corresponding reading text. A thorough analysis of the corpus and experimental results in this paper illustrate that our corpus ViMMRC demands reasoning abilities beyond simple word matching. We proposed the method of Boosted Sliding Window (BSW) that improves 5.51 human performance on the corpus and compared it to our MRC models. The performance gap between humans and our best experimental model indicates that significant progress can be made on Vietnamese machine reading comprehension in further research. The corpus is freely available at our website for research purposes.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/20/2020

An Experimental Study of Deep Neural Network Models for Vietnamese Multiple-Choice Reading Comprehension

Machine reading comprehension (MRC) is a challenging task in natural lan...
research
03/31/2023

A Multiple Choices Reading Comprehension Corpus for Vietnamese Language Education

Machine reading comprehension has been an interesting and challenging ta...
research
05/04/2021

Conversational Machine Reading Comprehension for Vietnamese Healthcare Texts

Machine reading comprehension (MRC) is a sub-field in natural language p...
research
12/20/2019

SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis

SberQuAD—a large scale analog of Stanford SQuAD in the Russian language—...
research
06/19/2020

New Vietnamese Corpus for Machine ReadingComprehension of Health News Articles

Although over 95 million people in the world speak the Vietnamese langua...
research
11/09/2019

Improving Machine Reading Comprehension via Adversarial Training

Adversarial training (AT) as a regularization method has proved its effe...
research
09/27/2019

Multi-Modal Citizen Science: From Disambiguation to Transcription of Classical Literature

The engagement of citizens in the research projects, including Digital H...

Please sign up or login with your details

Forgot password? Click here to reset