A Guide for Practical Use of ADMG Causal Data Augmentation

04/03/2023
by   Audrey Poinsot, et al.
0

Data augmentation is essential when applying Machine Learning in small-data regimes. It generates new samples following the observed data distribution while increasing their diversity and variability to help researchers and practitioners improve their models' robustness and, thus, deploy them in the real world. Nevertheless, its usage in tabular data still needs to be improved, as prior knowledge about the underlying data mechanism is seldom considered, limiting the fidelity and diversity of the generated data. Causal data augmentation strategies have been pointed out as a solution to handle these challenges by relying on conditional independence encoded in a causal graph. In this context, this paper experimentally analyzed the ADMG causal augmentation method considering different settings to support researchers and practitioners in understanding under which conditions prior knowledge helps generate new data points and, consequently, enhances the robustness of their models. The results highlighted that the studied method (a) is independent of the underlying model mechanism, (b) requires a minimal number of observations that may be challenging in a small-data regime to improve an ML model's accuracy, (c) propagates outliers to the augmented set degrading the performance of the model, and (d) is sensitive to its hyperparameter's value.

READ FULL TEXT
research
02/27/2021

Incorporating Causal Graphical Prior Knowledge into Predictive Modeling via Simple Data Augmentation

Causal graphs (CGs) are compact representations of the knowledge of the ...
research
01/26/2023

Experimenting with an Evaluation Framework for Imbalanced Data Learning (EFIDL)

Introduction Data imbalance is one of the crucial issues in big data ana...
research
03/03/2023

Unproportional mosaicing

Data shift is a gap between data distribution used for training and data...
research
02/09/2021

Negative Data Augmentation

Data augmentation is often used to enlarge datasets with synthetic sampl...
research
02/20/2020

Affinity and Diversity: Quantifying Mechanisms of Data Augmentation

Though data augmentation has become a standard component of deep neural ...
research
01/07/2022

GenLabel: Mixup Relabeling using Generative Models

Mixup is a data augmentation method that generates new data points by mi...
research
08/12/2023

DFM-X: Augmentation by Leveraging Prior Knowledge of Shortcut Learning

Neural networks are prone to learn easy solutions from superficial stati...

Please sign up or login with your details

Forgot password? Click here to reset