Around the GLOBE: Numerical Aggregation Question-Answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks

07/30/2023
by   Omri Suissa, et al.
0

One of the key AI tools for textual corpora exploration is natural language question-answering (QA). Unlike keyword-based search engines, QA algorithms receive and process natural language questions and produce precise answers to these questions, rather than long lists of documents that need to be manually scanned by the users. State-of-the-art QA algorithms based on DNNs were successfully employed in various domains. However, QA in the genealogical domain is still underexplored, while researchers in this field (and other fields in humanities and social sciences) can highly benefit from the ability to ask questions in natural language, receive concrete answers and gain insights hidden within large corpora. While some research has been recently conducted for factual QA in the genealogical domain, to the best of our knowledge, there is no previous research on the more challenging task of numerical aggregation QA (i.e., answering questions combining aggregation functions, e.g., count, average, max). Numerical aggregation QA is critical for distant reading and analysis for researchers (and the general public) interested in investigating cultural heritage domains. Therefore, in this study, we present a new end-to-end methodology for numerical aggregation QA for genealogical trees that includes: 1) an automatic method for training dataset generation; 2) a transformer-based table selection method, and 3) an optimized transformer-based numerical aggregation QA model. The findings indicate that the proposed architecture, GLOBE, outperforms the state-of-the-art models and pipelines by achieving 87 current state-of-the-art models. This study may have practical implications for genealogical information centers and museums, making genealogical data research easy and scalable for experts as well as the general public.

READ FULL TEXT
research
08/19/2021

UNIQORN: Unified Question Answering over RDF Knowledge Graphs and Natural Language Text

Question answering over knowledge graphs and other RDF data has been gre...
research
02/04/2022

Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Current research in natural language processing is highly dependent on c...
research
09/11/2018

How much should you ask? On the question structure in QA systems

Datasets that boosted state-of-the-art solutions for Question Answering ...
research
07/30/2023

Question Answering with Deep Neural Networks for Semi-Structured Heterogeneous Genealogical Knowledge Graphs

With the rising popularity of user-generated genealogical family trees, ...
research
08/08/2023

On Monotonic Aggregation for Open-domain QA

Question answering (QA) is a critical task for speech-based retrieval fr...
research
11/10/2021

Recent Advances in Automated Question Answering In Biomedical Domain

The objective of automated Question Answering (QA) systems is to provide...
research
03/01/2023

A Universal Question-Answering Platform for Knowledge Graphs

Knowledge from diverse application domains is organized as knowledge gra...

Please sign up or login with your details

Forgot password? Click here to reset