Evaluation of Tree Based Regression over Multiple Linear Regression for Non-normally Distributed Data in Battery Performance

11/03/2021
by   Shovan Chowdhury, et al.
0

Battery performance datasets are typically non-normal and multicollinear. Extrapolating such datasets for model predictions needs attention to such characteristics. This study explores the impact of data normality in building machine learning models. In this work, tree-based regression models and multiple linear regressions models are each built from a highly skewed non-normal dataset with multicollinearity and compared. Several techniques are necessary, such as data transformation, to achieve a good multiple linear regression model with this dataset; the most useful techniques are discussed. With these techniques, the best multiple linear regression model achieved an R^2 = 81.23 this study. Tree-based models perform better on this dataset, as they are non-parametric, capable of handling complex relationships among variables and not affected by multicollinearity. We show that bagging, in the use of Random Forests, reduces overfitting. Our best tree-based model achieved accuracy of R^2 = 97.73 machine learning model for non-normally distributed, multicollinear data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/19/2023

A New Bootstrap Goodness-of-Fit Test for Normal Linear Regression Models

In this work, the distributional properties of the goodness-of-fit term ...
research
10/16/2018

Hunting for Discriminatory Proxies in Linear Regression Models

A machine learning model may exhibit discrimination when used to make de...
research
07/29/2021

Deciphering Cryptic Behavior in Bimetallic Transition Metal Complexes with Machine Learning

The rational tailoring of transition metal complexes is necessary to add...
research
08/23/2023

Finding the Perfect Fit: Applying Regression Models to ClimateBench v1.0

Climate projections using data driven machine learning models acting as ...
research
08/13/2023

Optimizing Offensive Gameplan in the National Basketball Association with Machine Learning

Throughout the analytical revolution that has occurred in the NBA, the d...
research
02/03/2022

Machine Learning Solar Wind Driving Magnetospheric Convection in Tail Lobes

To quantitatively study the driving mechanisms of magnetospheric convect...
research
06/26/2017

Top-down Transformation Choice

Simple models are preferred over complex models, but over-simplistic mod...

Please sign up or login with your details

Forgot password? Click here to reset