Active Learning Methods based on Statistical Leverage Scores

12/06/2018
by   Cem Orhan, et al.
0

In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion that is based on statistical leverage scores and present two novel active learning methods based on this criterion: ALEVS for querying single example at each iteration and DBALEVS for querying a batch of examples. To assess the representativeness of the examples in the pool, ALEVS and DBALEVS use the statistical leverage scores of the kernel matrices computed on the examples of each class. Additionally, DBALEVS selects a diverse a set of examples that are highly representative but are dissimilar to already labeled examples through maximizing a submodular set function defined with the statistical leverage scores and the kernel matrix computed on the pool of the examples. The submodularity property of the set scoring function let us identify batches with a constant factor approximate to the optimal batch in an efficient manner. Our experiments on diverse datasets show that querying based on leverage scores is a powerful strategy for active learning.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/18/2019

The Effectiveness of Variational Autoencoders for Active Learning

The high cost of acquiring labels is one of the main challenges in deplo...
research
04/16/2021

Data Shapley Valuation for Efficient Batch Active Learning

Annotating the right set of data amongst all available data points is a ...
research
04/02/2019

Sequential Adaptive Design for Jump Regression Estimation

Selecting input data or design points for statistical models has been of...
research
06/27/2012

Batch Active Learning via Coordinated Matching

Most prior work on active learning of classifiers has focused on sequent...
research
06/30/2020

Similarity Search for Efficient Active Learning and Search of Rare Concepts

Many active learning and search approaches are intractable for industria...
research
06/23/2017

A Variance Maximization Criterion for Active Learning

Active learning aims to train a classifier as fast as possible with as f...
research
05/23/2017

Bayesian Pool-based Active Learning With Abstention Feedbacks

We study pool-based active learning with abstention feedbacks, where a l...

Please sign up or login with your details

Forgot password? Click here to reset