Active learning for reducing labeling effort in text classification tasks

09/10/2021
by   Pieter Floris Jacobs, et al.
0

Labeling data can be an expensive task as it is usually performed manually by domain experts. This is cumbersome for deep learning, as it is dependent on large labeled datasets. Active learning (AL) is a paradigm that aims to reduce labeling effort by only using the data which the used model deems most informative. Little research has been done on AL in a text classification setting and next to none has involved the more recent, state-of-the-art NLP models. Here, we present an empirical study that compares different uncertainty-based algorithms with BERT_base as the used classifier. We evaluate the algorithms on two NLP classification datasets: Stanford Sentiment Treebank and KvK-Frontpages. Additionally, we explore heuristics that aim to solve presupposed problems of uncertainty-based AL; namely, that it is unscalable and that it is prone to selecting outliers. Furthermore, we explore the influence of the query-pool size on the performance of AL. Whereas it was found that the proposed heuristics for AL did not improve performance of AL; our results show that using uncertainty-based AL with BERT_base outperforms random sampling of data. This difference in performance can decrease as the query-pool size gets larger.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
08/17/2020

A Survey of Active Learning for Text Classification using Deep Neural Networks

Natural language processing (NLP) and neural networks (NNs) have both un...
research
05/07/2022

Towards Computationally Feasible Deep Active Learning

Active learning (AL) is a prominent technique for reducing the annotatio...
research
07/09/2020

IALE: Imitating Active Learner Ensembles

Active learning (AL) prioritizes the labeling of the most informative da...
research
06/14/2016

Active Discriminative Text Representation Learning

We propose a new active learning (AL) method for text classification wit...
research
04/02/2020

In Automation We Trust: Investigating the Role of Uncertainty in Active Learning Systems

We investigate how different active learning (AL) query policies coupled...
research
09/02/2020

ALEX: Active Learning based Enhancement of a Model's Explainability

An active learning (AL) algorithm seeks to construct an effective classi...
research
11/30/2020

On Initial Pools for Deep Active Learning

Active Learning (AL) techniques aim to minimize the training data requir...

Please sign up or login with your details

Forgot password? Click here to reset