Building Interpretable and Reliable Open Information Retriever for New Domains Overnight

08/09/2023
by   Xiaodong Yu, et al.
0

Information retrieval (IR) or knowledge retrieval, is a critical component for many down-stream tasks such as open-domain question answering (QA). It is also very challenging, as it requires succinctness, completeness, and correctness. In recent works, dense retrieval models have achieved state-of-the-art (SOTA) performance on in-domain IR and QA benchmarks by representing queries and knowledge passages with dense vectors and learning the lexical and semantic similarity. However, using single dense vectors and end-to-end supervision are not always optimal because queries may require attention to multiple aspects and event implicit knowledge. In this work, we propose an information retrieval pipeline that uses entity/event linking model and query decomposition model to focus more accurately on different information units of the query. We show that, while being more interpretable and reliable, our proposed pipeline significantly improves passage coverages and denotation accuracies across five IR and QA benchmarks. It will be the go-to system to use for applications that need to perform IR on a new domain without much dedicated effort, because of its superior interpretability and cross-domain performance.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/02/2020

End-to-End QA on COVID-19: Domain Adaptation with Synthetic Training

End-to-end question answering (QA) requires both information retrieval (...
research
10/30/2019

Lexical Learning as an Online Optimal Experiment: Building Efficient Search Engines through Human-Machine Collaboration

Information retrieval (IR) systems need to constantly update their knowl...
research
04/30/2020

Progressively Pretrained Dense Corpus Index for Open-Domain Question Answering

To extract answers from a large corpus, open-domain question answering (...
research
02/14/2023

Large-Scale Knowledge Synthesis and Complex Information Retrieval from Biomedical Documents

Recent advances in the healthcare industry have led to an abundance of u...
research
04/23/2020

TCNN: Triple Convolutional Neural Network Models for Retrieval-based Question Answering System in E-commerce

Automatic question-answering (QA) systems have boomed during last few ye...
research
09/22/2020

Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering

Evidence retrieval is a key component of explainable question answering ...
research
04/18/2021

Simple and Efficient ways to Improve REALM

Dense retrieval has been shown to be effective for retrieving relevant d...

Please sign up or login with your details

Forgot password? Click here to reset