DeepAI AI Chat
Log In Sign Up

Testing for Regression Heteroskedasticity with High-Dimensional Random Forests

12/05/2022
by   Chi Chien-Ming, et al.
University of Florida
0

Statistical inference for high-dimensional regression heteroskedasticity is an important but under-explored problem. The current paper aims at filling this gap by proposing two tests, namely the variance difference test and the variance difference Breusch-Pagan test, for assessing high-dimensional regression heteroskedasticity. The former tests whether an explanatory feature of interest is associated with the conditional variance of a response variable, while the latter tests heteroskedasticity in the regression, which is known to be the Breusch-Pagan test problem. To formally establish the tests, we have derived rigorous P-values and test sizes, and analyzed the test power under a nonparametric heteroskedastic data generating model with high-dimensional input features. Such a model setting takes into account high-dimensional applications with flexible structures of heteroskedasticity and features having interaction effects on the mean of the response; these are common applications in many fields such as biology. Our methods leverage machine learning mean prediction methods such as random forests and use knockoff variables as negative controls. Particularly, the definition of knockoffs for our test statistics is more flexible than the original definition of knockoffs, and we give a detailed comparison of these two definitions and discuss the advantages of our knockoffs. The satisfactory empirical performance of the proposed tests is illustrated with simulation results and an HIV (Human Immunodeficiency Virus) case study.

READ FULL TEXT

page 1

page 2

page 3

page 4

12/27/2018

Power Comparison between High Dimensional t-Test, Sign, and Signed Rank Tests

In this paper, we propose a power comparison between high dimensional t-...
07/04/2022

FACT: High-Dimensional Random Forests Inference

Random forests is one of the most widely used machine learning methods o...
08/04/2015

Adaptivity and Computation-Statistics Tradeoffs for Kernel and Distance based High Dimensional Two Sample Testing

Nonparametric two sample testing is a decision theoretic problem that in...
01/31/2018

A Distribution-Free Test of Independence and Its Application to Variable Selection

Motivated by the importance of measuring the association between the res...
02/12/2019

Statistical inference with F-statistics when fitting simple models to high-dimensional data

We study linear subset regression in the context of the high-dimensional...
05/13/2020

Exchangeability, Conformal Prediction, and Rank Tests

Conformal prediction has been a very popular method of distribution-free...