Trimming Stability Selection increases variable selection robustness

11/23/2021
by   Tino Werner, et al.
0

Contamination can severely distort an estimator unless the estimation procedure is suitably robust. This is a well-known issue and has been addressed in Robust Statistics, however, the relation of contamination and distorted variable selection has been rarely considered in literature. As for variable selection, many methods for sparse model selection have been proposed, including Stability Selection which is a meta-algorithm based on some variable selection algorithm in order to immunize against particular data configurations. We introduce the variable selection breakdown point that quantifies the number of cases resp. cells that have to be contaminated in order to let no relevant variable be detected. We show that particular outlier configurations can completely mislead model selection and argue why even cell-wise robust methods cannot fix this problem. We combine the variable selection breakdown point with resampling, resulting in the Stability Selection breakdown point that quantifies the robustness of Stability Selection. We propose a trimmed Stability Selection which only aggregates the models with the lowest in-sample losses so that, heuristically, models computed on heavily contaminated resamples should be trimmed away. We provide a short simulation study that reveals both the potential of our approach as well as the fragility of variable selection, even for an extremely small cell-wise contamination rate.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/03/2019

Robust Model Selection for Finite Mixture of Regression Models Through Trimming

In this article, we introduce a new variable selection technique through...
research
04/27/2020

Using reference models in variable selection

Variable selection, or more generally, model reduction is an important a...
research
07/29/2020

Robust variable selection for model-based learning in presence of adulteration

The problem of identifying the most discriminating features when perform...
research
04/26/2017

Pruning variable selection ensembles

In the context of variable selection, ensemble learning has gained incre...
research
09/19/2021

Uncertainty quantification for robust variable selection and multiple testing

We study the problem of identifying the set of active variables, termed ...
research
07/30/2020

Accuracy and stability of solar variable selection comparison under complicated dependence structures

In this paper we focus on the variable-selection peformance of solar on ...
research
12/25/2019

Confounder Selection via Support Intersection

Confounding matters in almost all observational studies that focus on ca...

Please sign up or login with your details

Forgot password? Click here to reset