A Conditional Randomization Test for Sparse Logistic Regression in High-Dimension

05/29/2022
by   Binh T. Nguyen, et al.
8

Identifying the relevant variables for a classification model with correct confidence levels is a central but difficult task in high-dimension. Despite the core role of sparse logistic regression in statistics and machine learning, it still lacks a good solution for accurate inference in the regime where the number of features p is as large as or larger than the number of samples n. Here, we tackle this problem by improving the Conditional Randomization Test (CRT). The original CRT algorithm shows promise as a way to output p-values while making few assumptions on the distribution of the test statistics. As it comes with a prohibitive computational cost even in mildly high-dimensional problems, faster solutions based on distillation have been proposed. Yet, they rely on unrealistic hypotheses and result in low-power solutions. To improve this, we propose CRT-logit, an algorithm that combines a variable-distillation step and a decorrelation step that takes into account the geometry of ℓ_1-penalized logistic regression problem. We provide a theoretical analysis of this procedure, and demonstrate its effectiveness on simulations, along with experiments on large-scale brain-imaging and genomics datasets.

READ FULL TEXT
research
03/08/2017

Sparse Quadratic Logistic Regression in Sub-quadratic Time

We consider support recovery in the quadratic logistic regression settin...
research
03/23/2021

SLOE: A Faster Method for Statistical Inference in High-Dimensional Logistic Regression

Logistic regression remains one of the most widely used tools in applied...
research
11/03/2014

Bayesian feature selection with strongly-regularizing priors maps to the Ising Model

Identifying small subsets of features that are relevant for prediction a...
research
07/19/2016

Information-theoretical label embeddings for large-scale image classification

We present a method for training multi-label, massively multi-class imag...
research
01/20/2022

Using Machine Learning to Test Causal Hypotheses in Conjoint Analysis

Conjoint analysis is a popular experimental design used to measure multi...
research
10/03/2015

Distributed Parameter Map-Reduce

This paper describes how to convert a machine learning problem into a se...
research
05/11/2018

Stochastic Approximation EM for Logistic Regression with Missing Values

Logistic regression is a common classification method in supervised lear...

Please sign up or login with your details

Forgot password? Click here to reset