Best Split Nodes for Regression Trees

06/24/2019
by   Jason M. Klusowski, et al.
3

Decision trees with binary splits are popularly constructed using Classification and Regression Trees (CART) methodology. For regression models, at each node of the tree, the data is divided into two daughter nodes according to a split point that maximizes the reduction in variance (impurity) along a particular variable. This paper develops bounds on the size of a terminal node formed from a sequence of optimal splits via the infinite sample CART sum of squares criterion. We use these bounds to derive an interesting connection between the bias of a regression tree and the mean decrease in impurity (MDI) measure of variable importance---a tool widely used for model interpretability---defined as the weighted sum of impurity reductions over all nonterminal nodes in the tree. In particular, we show that the size of a terminal subnode for a variable is small when the MDI for that variable is large. Finally, we apply these bounds to show consistency of Breiman's random forests over a class of regression functions. The context is surprisingly general and applies to a wide variety of multivariable data generating distributions and regression functions. The main technical tool is an exact characterization of the conditional probabilities of the daughter nodes arising from an optimal split, in terms of the partial dependence function and reduction in impurity.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
06/07/2020

Sparse learning with CART

Decision trees with binary splits are popularly constructed using Classi...
research
10/19/2022

Distributional Adaptive Soft Regression Trees

Random forests are an ensemble method relevant for many problems, such a...
research
11/15/2007

Variable importance in binary regression trees and forests

We characterize and study variable importance (VIMP) and pairwise variab...
research
08/23/2022

Regularized impurity reduction: Accurate decision trees with complexity guarantees

Decision trees are popular classification models, providing high accurac...
research
11/05/2020

Nonparametric Variable Screening with Optimal Decision Stumps

Decision trees and their ensembles are endowed with a rich set of diagno...
research
10/30/2020

Measure Inducing Classification and Regression Trees for Functional Data

We propose a tree-based algorithm for classification and regression prob...
research
06/16/2016

The Effect of Heteroscedasticity on Regression Trees

Regression trees are becoming increasingly popular as omnibus predicting...

Please sign up or login with your details

Forgot password? Click here to reset