Knockoff Boosted Tree for Model-Free Variable Selection

02/20/2020
by   Tao Jiang, et al.
0

In this article, we propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. Our method is inspired by the original knockoff method, where the differences between original and knockoff variables are used for variable selection with false discovery rate control. The original method uses Lasso for regression models and assumes there are more samples than variables. We extend this method to both model-free and high-dimensional variable selection. We propose two new sampling methods for generating knockoffs, namely the sparse covariance and principal component knockoff methods. We test these methods and compare them with the original knockoff method in terms of their ability to control type I errors and power. The boosted tree model is a complex system and has more hyperparameters than models with simpler assumptions. In our framework, these hyperparameters are either tuned through Bayesian optimization or fixed at multiple levels for trend detection. In simulation tests, we also compare the properties and performance of importance test statistics of tree models. The results include combinations of different knockoffs and importance test statistics. We also consider scenarios that include main-effect, interaction, exponential, and second-order models while assuming the true model structures are unknown. We apply our algorithm for tumor purity estimation and tumor classification using the Cancer Genome Atlas (TCGA) gene expression data. The proposed algorithm is included in the KOBT package, available at <https://cran.r-project.org/web/packages/KOBT/index.html>.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/28/2019

Variable Selection with Copula Entropy

Variable selection is of significant importance for classification and r...
research
07/17/2020

An Easy-to-Implement Hierarchical Standardization for Variable Selection Under Strong Heredity Constraint

For many practical problems, the regression models follow the strong her...
research
09/13/2019

SuRF: a New Method for Sparse Variable Selection, with Application in Microbiome Data Analysis

In this paper, we present a new variable selection method for regression...
research
11/16/2018

Deep Knockoffs

This paper introduces a machine for sampling approximate model-X knockof...
research
04/24/2023

High-dimensional iterative variable selection for accelerated failure time models

We propose an iterative variable selection method for the accelerated fa...
research
03/06/2022

Variable Selection with the Knockoffs: Composite Null Hypotheses

The Fixed-X knockoff filter is a flexible framework for variable selecti...
research
12/10/2018

Variational Nonparametric Discriminant Analysis

Variable selection and classification methods are common objectives in t...

Please sign up or login with your details

Forgot password? Click here to reset