Conditional Synthetic Data Generation for Robust Machine Learning Applications with Limited Pandemic Data

by   Hari Prasanna Das, et al.

Background: At the onset of a pandemic, such as COVID-19, data with proper labeling/attributes corresponding to the new disease might be unavailable or sparse. Machine Learning (ML) models trained with the available data, which is limited in quantity and poor in diversity, will often be biased and inaccurate. At the same time, ML algorithms designed to fight pandemics must have good performance and be developed in a time-sensitive manner. To tackle the challenges of limited data, and label scarcity in the available data, we propose generating conditional synthetic data, to be used alongside real data for developing robust ML models. Methods: We present a hybrid model consisting of a conditional generative flow and a classifier for conditional synthetic data generation. The classifier decouples the feature representation for the condition, which is fed to the flow to extract the local noise. We generate synthetic data by manipulating the local noise with fixed conditional feature representation. We also propose a semi-supervised approach to generate synthetic samples in the absence of labels for a majority of the available data. Results: We performed conditional synthetic generation for chest computed tomography (CT) scans corresponding to normal, COVID-19, and pneumonia afflicted patients. We show that our method significantly outperforms existing models both on qualitative and quantitative performance, and our semi-supervised approach can efficiently synthesize conditional samples under label scarcity. As an example of downstream use of synthetic data, we show improvement in COVID-19 detection from CT scans with conditional synthetic data augmentation.


page 2

page 3

page 4

page 5


Improving COVID-19 CXR Detection with Synthetic Data Augmentation

Since the beginning of the COVID-19 pandemic, researchers have developed...

On the Usefulness of Synthetic Tabular Data Generation

Despite recent advances in synthetic data generation, the scientific com...

3D Tomographic Pattern Synthesis for Enhancing the Quantification of COVID-19

The Coronavirus Disease (COVID-19) has affected 1.8 million people and r...

Copula-based synthetic data generation for machine learning emulators in weather and climate: application to a simple radiation model

Can we improve machine learning (ML) emulators with synthetic data? The ...

Synthetic flow-based cryptomining attack generation through Generative Adversarial Networks

Due to the growing rise of cyber attacks in the Internet, flow-based dat...

BinarySDG: binary sensor data generation with R

The scarcity of Smart Home data is still a pretty big problem, and in a ...

Integrating Expert ODEs into Neural ODEs: Pharmacology and Disease Progression

Modeling a system's temporal behaviour in reaction to external stimuli i...

Please sign up or login with your details

Forgot password? Click here to reset