The SKIM-FA Kernel: High-Dimensional Variable Selection and Nonlinear Interaction Discovery in Linear Time

06/23/2021
by   Raj Agrawal, et al.
0

Many scientific problems require identifying a small set of covariates that are associated with a target response and estimating their effects. Often, these effects are nonlinear and include interactions, so linear and additive methods can lead to poor estimation and variable selection. Unfortunately, methods that simultaneously express sparsity, nonlinearity, and interactions are computationally intractable – with runtime at least quadratic in the number of covariates, and often worse. In the present work, we solve this computational bottleneck. We show that suitable interaction models have a kernel representation, namely there exists a "kernel trick" to perform variable selection and estimation in O(# covariates) time. Our resulting fit corresponds to a sparse orthogonal decomposition of the regression function in a Hilbert space (i.e., a functional ANOVA decomposition), where interaction effects represent all variation that cannot be explained by lower-order effects. On a variety of synthetic and real datasets, our approach outperforms existing methods used for large, high-dimensional datasets while remaining competitive (or being orders of magnitude faster) in runtime.

READ FULL TEXT
research
04/19/2018

Large-scale Nonlinear Variable Selection via Kernel Random Features

We propose a new method for input variable selection in nonlinear regres...
research
06/03/2019

Semi-parametric Bayesian variable selection for gene-environment interactions

Many complex diseases are known to be affected by the interactions betwe...
research
02/26/2018

Scalable kernel-based variable selection with sparsistency

Variable selection is central to high-dimensional data analysis, and var...
research
11/12/2019

Purifying Interaction Effects with the Functional ANOVA: An Efficient Algorithm for Recovering Identifiable Additive Models

Recent methods for training generalized additive models (GAMs) with pair...
research
10/17/2016

The xyz algorithm for fast interaction search in high-dimensional data

When performing regression on a dataset with p variables, it is often of...
research
05/16/2019

The Kernel Interaction Trick: Fast Bayesian Discovery of Pairwise Interactions in High Dimensions

Discovering interaction effects on a response of interest is a fundament...

Please sign up or login with your details

Forgot password? Click here to reset