Convex space learning improves deep-generative oversampling for tabular imbalanced classification on smaller datasets

06/20/2022
by   Kristian Schultz, et al.
6

Data is commonly stored in tabular format. Several fields of research (e.g., biomedical, fault/fraud detection), are prone to small imbalanced tabular data. Supervised Machine Learning on such data is often difficult due to class imbalance, adding further to the challenge. Synthetic data generation i.e. oversampling is a common remedy used to improve classifier performance. State-of-the-art linear interpolation approaches, such as LoRAS and ProWRAS can be used to generate synthetic samples from the convex space of the minority class to improve classifier performance in such cases. Generative Adversarial Networks (GANs) are common deep learning approaches for synthetic sample generation. Although GANs are widely used for synthetic image generation, their scope on tabular data in the context of imbalanced classification is not adequately explored. In this article, we show that existing deep generative models perform poorly compared to linear interpolation approaches generating synthetic samples from the convex space of the minority class, for imbalanced classification problems on tabular datasets of small size. We propose a deep generative model, ConvGeN combining the idea of convex space learning and deep generative models. ConVGeN learns the coefficients for the convex combinations of the minority class samples, such that the synthetic data is distinct enough from the majority class. We demonstrate that our proposed model ConvGeN improves imbalanced classification on such small datasets, as compared to existing deep generative models while being at par with the existing linear interpolation approaches. Moreover, we discuss how our model can be used for synthetic tabular data generation in general, even outside the scope of data imbalance, and thus, improves the overall applicability of convex space learning.

READ FULL TEXT

page 5

page 6

page 9

page 10

page 11

page 14

page 15

page 19

research
05/07/2020

Minority Class Oversampling for Tabular Data with Deep Generative Models

In practice, data scientists are often confronted with imbalanced data. ...
research
01/03/2021

Synthetic Embedding-based Data Generation Methods for Student Performance

Given the inherent class imbalance issue within student performance data...
research
01/14/2022

Synthesising Electronic Health Records: Cystic Fibrosis Patient Group

Class imbalance can often degrade predictive performance of supervised l...
research
08/16/2023

Fair GANs through model rebalancing with synthetic data

Deep generative models require large amounts of training data. This ofte...
research
05/17/2021

Synthesising Multi-Modal Minority Samples for Tabular Data

Real-world binary classification tasks are in many cases imbalanced, whe...
research
06/16/2023

Understanding Deep Generative Models with Generalized Empirical Likelihoods

Understanding how well a deep generative model captures a distribution o...
research
04/29/2023

Analyzing drop coalescence in microfluidic device with a deep learning generative model

Predicting drop coalescence based on process parameters is crucial for e...

Please sign up or login with your details

Forgot password? Click here to reset