Bootstrapped Edge Count Tests for Nonparametric Two-Sample Inference Under Heterogeneity

04/26/2023
by   Trambak Banerjee, et al.
0

Nonparametric two-sample testing is a classical problem in inferential statistics. While modern two-sample tests, such as the edge count test and its variants, can handle multivariate and non-Euclidean data, contemporary gargantuan datasets often exhibit heterogeneity due to the presence of latent subpopulations. Direct application of these tests, without regulating for such heterogeneity, may lead to incorrect statistical decisions. We develop a new nonparametric testing procedure that accurately detects differences between the two samples in the presence of unknown heterogeneity in the data generation process. Our framework handles this latent heterogeneity through a composite null that entertains the possibility that the two samples arise from a mixture distribution with identical component distributions but with possibly different mixing weights. In this regime, we study the asymptotic behavior of weighted edge count test statistic and show that it can be effectively re-calibrated to detect arbitrary deviations from the composite null. For practical implementation we propose a Bootstrapped Weighted Edge Count test which involves a bootstrap-based calibration procedure that can be easily implemented across a wide range of heterogeneous regimes. A comprehensive simulation study and an application to detecting aberrant user behaviors in online games demonstrates the excellent non-asymptotic performance of the proposed test.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/05/2020

A Nearest-Neighbor Based Nonparametric Test for Viral Remodeling in Heterogeneous Single-Cell Proteomic Data

An important problem in contemporary immunology studies based on single-...
research
06/07/2023

Multivariate two-sample test statistics based on data depth

Data depth has been applied as a nonparametric measurement for ranking m...
research
12/24/2021

RISE: Rank in Similarity Graph Edge-Count Two-Sample Test

Two-sample hypothesis testing for high-dimensional data is ubiquitous no...
research
12/22/2021

Omnibus goodness-of-fit tests for count distributions

A consistent omnibus goodness-of-fit test for count distributions is pro...
research
06/13/2016

Tuning-Free Heterogeneity Pursuit in Massive Networks

Heterogeneity is often natural in many contemporary applications involvi...
research
05/04/2022

Validating Approximate Slope Homogeneity in Large Panels

Statistical inference for large data panels is omnipresent in modern eco...
research
07/06/2021

Testing for the Presence of Structural Change and Spatial Heterogeneity

In a spatial-temporal model, structural change and/or spatial heterogene...

Please sign up or login with your details

Forgot password? Click here to reset