An "outside the box" solution for imbalanced data classification

11/16/2019
by   Hubert Jegierski, et al.
86

A common problem of the real-world data sets is the class imbalance, which can significantly affect the classification abilities of classifiers. Numerous methods have been proposed to cope with this problem; however, even state-of-the-art methods offer a limited improvement (if any) for data sets with critically under-represented minority classes. For such problematic cases, an "outside the box" solution is required. Therefore, we propose a novel technique, called enrichment, which uses the information (observations) from the external data set(s). We present three approaches to implement enrichment technique: (1) selecting observations randomly, (2) iteratively choosing observations that improve the classification result, (3) adding observations that help the classifier to determine the border between classes better. We then thoroughly analyze developed solutions on ten real-world data sets to experimentally validate their usefulness. On average, our best approach improves the classification quality by 27%, and in the best case, by outstanding 66%. We also compare our technique with the universally applicable state-of-the-art methods. We find that our technique surpasses the existing methods performing, on average, 21% better. The advantage is especially noticeable for the smallest data sets, for which existing methods failed, while our solutions achieved the best results. Additionally, our technique applies to both the multi-class and binary classification tasks. It can also be combined with other techniques dealing with the class imbalance problem.

READ FULL TEXT

page 8

page 9

page 18

page 21

page 22

page 23

page 24

page 25

research
03/26/2016

A generalized flow for multi-class and binary classification tasks: An Azure ML approach

The constant growth in the present day real-world databases pose computa...
research
08/21/2023

Evaluating quantum generative models via imbalanced data classification benchmarks

A limited set of tools exist for assessing whether the behavior of quant...
research
04/07/2020

Combined Cleaning and Resampling Algorithm for Multi-Class Imbalanced Data with Label Noise

The imbalanced data classification is one of the most crucial tasks faci...
research
01/08/2018

Bridging the Gap: Simultaneous Fine Tuning for Data Re-Balancing

There are many real-world classification problems wherein the issue of d...
research
07/27/2023

Retrieval-based Text Selection for Addressing Class-Imbalanced Data in Classification

This paper addresses the problem of selecting of a set of texts for anno...
research
02/26/2016

Large-Scale Detection of Non-Technical Losses in Imbalanced Data Sets

Non-technical losses (NTL) such as electricity theft cause significant h...
research
07/12/2017

Influence of Resampling on Accuracy of Imbalanced Classification

In many real-world binary classification tasks (e.g. detection of certai...

Please sign up or login with your details

Forgot password? Click here to reset