VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

05/20/2023
by   Dao Xuan-Quy, et al.
0

The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article. The dataset, which covers nine subjects, was generated from the Vietnamese National High School Graduation Examination and comparable tests. 300 literary essays have been included, and there are over 19,000 multiple-choice questions on a range of topics. The dataset assesses LLMs in multitasking situations such as question answering, text generation, reading comprehension, visual question answering, and more by including both textual data and accompanying images. Using ChatGPT and BingChat, we evaluated LLMs on the VNHSGE dataset and contrasted their performance with that of Vietnamese students to see how well they performed. The results show that ChatGPT and BingChat both perform at a human level in a number of areas, including literature, English, history, geography, and civics education. They still have space to grow, though, especially in the areas of mathematics, physics, chemistry, and biology. The VNHSGE dataset seeks to provide an adequate benchmark for assessing the abilities of LLMs with its wide-ranging coverage and variety of activities. We intend to promote future developments in the creation of LLMs by making this dataset available to the scientific community, especially in resolving LLMs' limits in disciplines involving mathematics and the natural sciences.

READ FULL TEXT

page 15

page 16

page 17

page 18

page 19

research
05/03/2023

NorQuAD: Norwegian Question Answering Dataset

In this paper we present NorQuAD: the first Norwegian question answering...
research
07/17/2023

ChatGPT is Good but Bing Chat is Better for Vietnamese Students

This study examines the efficacy of two SOTA large language models (LLMs...
research
09/27/2021

FQuAD2.0: French Question Answering and knowing that you know nothing

Question Answering, including Reading Comprehension, is one of the NLP r...
research
09/19/2023

Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change

Pirá is a reading comprehension dataset focused on the ocean, the Brazil...
research
01/31/2023

Mathematical Capabilities of ChatGPT

We investigate the mathematical capabilities of ChatGPT by testing it on...
research
04/28/2023

ChatGPT – a Blessing or a Curse for Undergraduate Computer Science Students and Instructors?

ChatGPT is an AI language model developed by OpenAI that can understand ...

Please sign up or login with your details

Forgot password? Click here to reset