Differentially Private Synthetic Data: Applied Evaluations and Enhancements

11/11/2020
by   Lucas Rosenblatt, et al.
0

Machine learning practitioners frequently seek to leverage the most informative available data, without violating the data owner's privacy, when building predictive models. Differentially private data synthesis protects personal details from exposure, and allows for the training of differentially private machine learning models on privately generated datasets. But how can we effectively assess the efficacy of differentially private synthetic data? In this paper, we survey four differentially private generative adversarial networks for data synthesis. We evaluate each of them at scale on five standard tabular datasets, and in two applied industry scenarios. We benchmark with novel metrics from recent literature and other standard machine learning tools. Our results suggest some synthesizers are more applicable for different privacy budgets, and we further demonstrate complicating domain-based tradeoffs in selecting an approach. We offer experimental learning on applied machine learning scenarios with private internal data to researchers and practioners alike. In addition, we propose QUAIL, an ensemble-based modeling approach to generating synthetic data. We examine QUAIL's tradeoffs, and note circumstances in which it outperforms baseline differentially private supervised learning models under the same budget constraint.

READ FULL TEXT
research
11/28/2019

Comparative Study of Differentially Private Synthetic Data Algorithms and Evaluation Standards

Differentially private synthetic data generation is becoming a popular s...
research
05/04/2023

Leveraging gradient-derived metrics for data selection and valuation in differentially private training

Obtaining high-quality data for collaborative training of machine learni...
research
06/15/2020

GS-WGAN: A Gradient-Sanitized Approach for Learning Differentially Private Generators

The wide-spread availability of rich data has fueled the growth of machi...
research
04/20/2023

DPAF: Image Synthesis via Differentially Private Aggregation in Forward Phase

Differentially private synthetic data is a promising alternative for sen...
research
12/22/2020

Differentially Private Synthetic Medical Data Generation using Convolutional GANs

Deep learning models have demonstrated superior performance in several a...
research
12/09/2021

Differentially Private Ensemble Classifiers for Data Streams

Learning from continuous data streams via classification/regression is p...
research
08/08/2023

Accurate, Explainable, and Private Models: Providing Recourse While Minimizing Training Data Leakage

Machine learning models are increasingly utilized across impactful domai...

Please sign up or login with your details

Forgot password? Click here to reset