A Differentially Private Kernel Two-Sample Test

08/01/2018
by   Anant Raj, et al.
2

Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples can expose sensitive information about individuals who participate in scientific studies, which makes the current tests vulnerable to privacy breaches. Hence, we design a new framework for kernel two-sample testing conforming to differential privacy constraints, in order to guarantee the privacy of subjects in the data. Unlike existing differentially private parametric tests that simply add noise to data, kernel-based testing imposes a challenge due to a complex dependence of test statistics on the raw data, as these statistics correspond to estimators of distances between representations of probability measures in Hilbert spaces. Our approach considers finite dimensional approximations to those representations. As a result, a simple chi-squared test is obtained, where a test statistic depends on a mean and covariance of empirical differences between the samples, which we perturb for a privacy guarantee. We investigate the utility of our framework in two realistic settings and conclude that our method requires only a relatively modest increase in sample size to achieve a similar level of power to the non-private tests in both settings.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/22/2019

Differentially Private Nonparametric Hypothesis Testing

Hypothesis tests are a crucial statistical tool for data mining and are ...
research
06/24/2021

Covariance-Aware Private Mean Estimation Without Private Covariance Estimation

We present two sample-efficient differentially private mean estimators f...
research
07/18/2017

Differentially Private Identity and Closeness Testing of Discrete Distributions

We investigate the problems of identity and closeness testing over a dis...
research
10/14/2019

Two-sample Testing Using Deep Learning

We propose a two-sample testing procedure based on learned deep neural n...
research
11/03/2017

Differentially Private ANOVA Testing

Modern society generates an incredible amount of data about individuals,...
research
05/29/2023

Unleashing the Power of Randomization in Auditing Differentially Private ML

We present a rigorous methodology for auditing differentially private ma...
research
09/20/2021

The power of private likelihood-ratio tests for goodness-of-fit in frequency tables

Privacy-protecting data analysis investigates statistical methods under ...

Please sign up or login with your details

Forgot password? Click here to reset