Flexible variable selection in the presence of missing data

02/25/2022
by   B. D. Williamson, et al.
0

In many applications, it is of interest to identify a parsimonious set of features, or panel, from multiple candidates that achieves a desired level of performance in predicting a response. This task is often complicated in practice by missing data arising from the sampling design or other random mechanisms. Most recent work on variable selection in missing data contexts relies in some part on a finite-dimensional statistical model (e.g., a generalized or penalized linear model). In cases where this model is misspecified, the selected variables may not all be truly scientifically relevant and can result in panels with suboptimal classification performance. To address this limitation, we propose several nonparametric variable selection algorithms combined with multiple imputation to develop flexible panels in the presence of missing-at-random data. We outline strategies based on the proposed algorithms that achieve control of commonly used error rates. Through simulations, we show that our proposals have good operating characteristics and result in panels with higher classification performance compared to several existing penalized regression approaches. Finally, we use the proposed methods to develop biomarker panels for separating pancreatic cysts with differing malignancy potential in a setting where complicated missingness in the biomarkers arose due to limited specimen volumes.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/06/2021

Variable selection with missing data in both covariates and outcomes: Imputation and machine learning

The missing data issue is ubiquitous in health studies. Variable selecti...
research
02/26/2022

Missing Value Knockoffs

One limitation of the most statistical/machine learning-based variable s...
research
11/10/2021

variable selection and missing data imputation in categorical genomic data analysis by integrated ridge regression and random forest

Genomic data arising from a genome-wide association study (GWAS) are oft...
research
03/16/2020

Variable selection with multiply-imputed datasets: choosing between stacked and grouped methods

Penalized regression methods, such as lasso and elastic net, are used in...
research
08/16/2022

Variable Selection in Latent Regression IRT Models via Knockoffs: An Application to International Large-scale Assessment in Education

International large-scale assessments (ILSAs) play an important role in ...
research
04/04/2018

Variable selection using pseudo-variables

Penalized regression has become a standard tool for model building acros...
research
01/14/2019

Supervised Learning for Multi-Block Incomplete Data

In the supervised high dimensional settings with a large number of varia...

Please sign up or login with your details

Forgot password? Click here to reset