Optimal subgroup selection

09/02/2021
by   Henry W J Reeve, et al.
22

In clinical trials and other applications, we often see regions of the feature space that appear to exhibit interesting behaviour, but it is unclear whether these observed phenomena are reflected at the population level. Focusing on a regression setting, we consider the subgroup selection challenge of identifying a region of the feature space on which the regression function exceeds a pre-determined threshold. We formulate the problem as one of constrained optimisation, where we seek a low-complexity, data-dependent selection set on which, with a guaranteed probability, the regression function is uniformly at least as large as the threshold; subject to this constraint, we would like the region to contain as much mass under the marginal feature distribution as possible. This leads to a natural notion of regret, and our main contribution is to determine the minimax optimal rate for this regret in both the sample size and the Type I error probability. The rate involves a delicate interplay between parameters that control the smoothness of the regression function, as well as exponents that quantify the extent to which the optimal selection set at the population level can be approximated by families of well-behaved subsets. Finally, we expand the scope of our previous results by illustrating how they may be generalised to a treatment and control setting, where interest lies in the heterogeneous treatment effect.

READ FULL TEXT
POST COMMENT

Comments

There are no comments yet.

Authors

page 15

page 16

02/14/2019

Classification with unknown class conditional label noise on non-compact feature spaces

We investigate the problem of classification in the presence of unknown ...
12/13/2019

Assessing effect heterogeneity of a randomized treatment using conditional inference trees

Treatment effect heterogeneity occurs when individual characteristics in...
07/06/2020

On optimal two-stage testing of multiple mediators

Mediation analysis in high-dimensional settings often involves identifyi...
09/08/2020

Designing Transportable Experiments

We consider the problem of designing a randomized experiment on a source...
08/14/2021

Evidence Aggregation for Treatment Choice

Consider a planner who has to decide whether or not to introduce a new p...
04/09/2019

Evaluating Competence Measures for Dynamic Regressor Selection

Dynamic regressor selection (DRS) systems work by selecting the most com...
02/14/2022

Domain-Adjusted Regression or: ERM May Already Learn Features Sufficient for Out-of-Distribution Generalization

A common explanation for the failure of deep networks to generalize out-...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.