Active Sequential Two-Sample Testing

01/30/2023
by   Weizhi Li, et al.
0

Two-sample testing tests whether the distributions generating two samples are identical. We pose the two-sample testing problem in a new scenario where the sample measurements (or sample features) are inexpensive to access, but their group memberships (or labels) are costly. We devise the first active sequential two-sample testing framework that not only sequentially but also actively queries sample labels to address the problem. Our test statistic is a likelihood ratio where one likelihood is found by maximization over all class priors, and the other is given by a classification model. The classification model is adaptively updated and then used to guide an active query scheme called bimodal query to label sample features in the regions with high dependency between the feature variables and the label variables. The theoretical contributions in the paper include proof that our framework produces an anytime-valid p-value; and, under reachable conditions and a mild assumption, the framework asymptotically generates a minimum normalized log-likelihood ratio statistic that a passive query scheme can only achieve when the feature variable and the label variable have the highest dependence. Lastly, we provide a query-switching (QS) algorithm to decide when to switch from passive query to active query and adapt bimodal query to increase the testing power of our test. Extensive experiments justify our theoretical contributions and the effectiveness of QS.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/17/2021

A label efficient two-sample test

Two-sample tests evaluate whether two samples are realizations of the sa...
research
03/09/2021

On testing mean proportionality of multivariate normal variables

This short note considers the problem of testing the null hypothesis tha...
research
01/23/2019

kd-switch: A Universal Online Predictor with an application to Sequential Two-Sample Testing

We propose a novel online predictor for discrete labels conditioned on m...
research
01/16/2014

A Model-Based Active Testing Approach to Sequential Diagnosis

Model-based diagnostic reasoning often leads to a large number of diagno...
research
05/27/2016

Asymptotic Analysis of Objectives based on Fisher Information in Active Learning

Obtaining labels can be costly and time-consuming. Active learning allow...
research
09/05/2022

GRASP: A Goodness-of-Fit Test for Classification Learning

Performance of classifiers is often measured in terms of average accurac...
research
07/03/2020

Two-sample Testing for Large, Sparse High-Dimensional Multinomials under Rare/Weak Perturbations

Given two samples from possibly different discrete distributions over a ...

Please sign up or login with your details

Forgot password? Click here to reset