Scalable Evaluation and Improvement of Document Set Expansion via Neural Positive-Unlabeled Learning

10/29/2019
by   Alon Jacovi, et al.
0

We consider the situation in which a user has collected a small set of documents on a cohesive topic, and they want to retrieve additional documents on this topic from a large collection. Information Retrieval (IR) solutions treat the document set as a query, and look for similar documents in the collection. We propose to extend the IR approach by treating the problem as an instance of positive-unlabeled (PU) learning—i.e., learning binary classifiers from only positive and unlabeled data, where the positive data corresponds to the query documents, and the unlabeled data is the results returned by the IR engine. Utilizing PU learning for text with big neural networks is a largely unexplored field. We discuss various challenges in applying PU learning to the setting, including an unknown class prior, extremely imbalanced data and large-scale accurate evaluation of models, and we propose solutions and empirically validate them. We demonstrate the effectiveness of the method using a series of experiments of retrieving PubMed abstracts adhering to fine-grained topics. We demonstrate improvements over the base IR solution and other baselines. Implementation is available at https://github.com/sayaendo/document-set-expansion-pu.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/03/2021

Unsupervised Document Expansion for Information Retrieval with Stochastic Text Generation

One of the challenges in information retrieval (IR) is the vocabulary mi...
research
12/10/2019

Neural-IR-Explorer: A Content-Focused Tool to Explore Neural Re-Ranking Results

In this paper we look beyond metrics-based evaluation of Information Ret...
research
01/30/2013

Query Expansion in Information Retrieval Systems using a Bayesian Network-Based Thesaurus

Information Retrieval (IR) is concerned with the identification of docum...
research
12/27/2020

Neural document expansion for ad-hoc information retrieval

Recently, Nogueira et al. [2019] proposed a new approach to document exp...
research
08/29/2023

Class Prior-Free Positive-Unlabeled Learning with Taylor Variational Loss for Hyperspectral Remote Sensing Imagery

Positive-unlabeled learning (PU learning) in hyperspectral remote sensin...
research
10/05/2010

A bagging SVM to learn from positive and unlabeled examples

We consider the problem of learning a binary classifier from a training ...
research
04/07/2023

T2Ranking: A large-scale Chinese Benchmark for Passage Ranking

Passage ranking involves two stages: passage retrieval and passage re-ra...

Please sign up or login with your details

Forgot password? Click here to reset