Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

02/04/2022
by   André F. A. Paschoal, et al.
0

Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pirá dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pirá is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pirá dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pirá, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.

READ FULL TEXT
research
03/06/2023

AmQA: Amharic Question Answering Dataset

Question Answering (QA) returns concise answers or answer lists from nat...
research
11/24/2022

Question Answering and Question Generation for Finnish

Recent advances in the field of language modeling have improved the stat...
research
05/04/2022

KenSwQuAD – A Question Answering Dataset for Swahili Low Resource Language

This research developed a Kencorpus Swahili Question Answering Dataset K...
research
07/30/2023

Around the GLOBE: Numerical Aggregation Question-Answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks

One of the key AI tools for textual corpora exploration is natural langu...
research
04/24/2020

Question Answering over Curated and Open Web Sources

The last few years have seen an explosion of research on the topic of au...
research
01/01/2021

NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned

We review the EfficientQA competition from NeurIPS 2020. The competition...
research
09/28/2020

What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Open domain question answering (OpenQA) tasks have been recently attract...

Please sign up or login with your details

Forgot password? Click here to reset