Is BERT Robust to Label Noise? A Study on Learning with Noisy Labels in Text Classification

04/20/2022
by   Dawei Zhu, et al.
0

Incorrect labels in training data occur when human annotators make mistakes or when the data is generated via weak or distant supervision. It has been shown that complex noise-handling techniques - by modeling, cleaning or filtering the noisy instances - are required to prevent models from fitting this label noise. However, we show in this work that, for text classification tasks with modern NLP models like BERT, over a variety of noise types, existing noisehandling methods do not always improve its performance, and may even deteriorate it, suggesting the need for further investigation. We also back our observations with a comprehensive analysis.

READ FULL TEXT

page 4

page 7

research
01/27/2021

Towards Robustness to Label Noise in Text Classification via Noise Modeling

Large datasets in NLP suffer from noisy labels, due to erroneous automat...
research
06/03/2022

Task-Adaptive Pre-Training for Boosting Learning With Noisy Labels: A Study on Text Classification for African Languages

For high-resource languages like English, text classification is a well-...
research
01/24/2021

Analysing the Noise Model Error for Realistic Noisy Label Data

Distant and weak supervision allow to obtain large amounts of labeled tr...
research
05/01/2020

Learning from Noisy Labels with Noise Modeling Network

Multi-label image classification has generated significant interest in r...
research
09/11/2018

Training and Prediction Data Discrepancies: Challenges of Text Classification with Noisy, Historical Data

Industry datasets used for text classification are rarely created for th...
research
11/14/2017

On Extending Neural Networks with Loss Ensembles for Text Classification

Ensemble techniques are powerful approaches that combine several weak le...
research
12/30/2022

Distant Reading of the German Coalition Deal: Recognizing Policy Positions with BERT-based Text Classification

Automated text analysis has become a widely used tool in political scien...

Please sign up or login with your details

Forgot password? Click here to reset