A Pipeline for Generating, Annotating and Employing Synthetic Data for Real World Question Answering

11/30/2022
by   Matthew Maufe, et al.
0

Question Answering (QA) is a growing area of research, often used to facilitate the extraction of information from within documents. State-of-the-art QA models are usually pre-trained on domain-general corpora like Wikipedia and thus tend to struggle on out-of-domain documents without fine-tuning. We demonstrate that synthetic domain-specific datasets can be generated easily using domain-general models, while still providing significant improvements to QA performance. We present two new tools for this task: A flexible pipeline for validating the synthetic QA data and training downstream models on it, and an online interface to facilitate human annotation of this generated data. Using this interface, crowdworkers labelled 1117 synthetic QA pairs, which we then used to fine-tune downstream models and improve domain-specific QA performance by 8.75 F1.

READ FULL TEXT
research
11/24/2022

Question Answering and Question Generation for Finnish

Recent advances in the field of language modeling have improved the stat...
research
04/02/2018

Simple and Effective Semi-Supervised Question Answering

Recent success of deep learning models for the task of extractive Questi...
research
06/12/2019

Synthetic QA Corpora Generation with Roundtrip Consistency

We introduce a novel method of generating synthetic question answering c...
research
04/07/2022

Parameter-Efficient Abstractive Question Answering over Tables or Text

A long-term ambition of information seeking QA systems is to reason over...
research
04/17/2022

WikiOmnia: generative QA corpus on the whole Russian Wikipedia

The General QA field has been developing the methodology referencing the...
research
05/14/2023

Learning to Generalize for Cross-domain QA

There have been growing concerns regarding the out-of-domain generalizat...
research
09/12/2019

Measuring Domain Portability and ErrorPropagation in Biomedical QA

In this work we present Google's submission to the BioASQ 7 biomedical q...

Please sign up or login with your details

Forgot password? Click here to reset