Model Selection and Overfitting in Genetic Programming: Empirical Study [Extended Version]

04/30/2015
by   Jan Žegklitz, et al.
0

Genetic Programming has been very successful in solving a large area of problems but its use as a machine learning algorithm has been limited so far. One of the reasons is the problem of overfitting which cannot be solved or suppresed as easily as in more traditional approaches. Another problem, closely related to overfitting, is the selection of the final model from the population. In this article we present our research that addresses both problems: overfitting and model selection. We compare several ways of dealing with ovefitting, based on Random Sampling Technique (RST) and on using a validation set, all with an emphasis on model selection. We subject each approach to a thorough testing on artificial and real--world datasets and compare them with the standard approach, which uses the full training data, as a baseline.

READ FULL TEXT
research
02/16/2018

Train on Validation: Squeezing the Data Lemon

Model selection on validation data is an essential step in machine learn...
research
12/06/2017

On overfitting and post-selection uncertainty assessments

In a regression context, when the relevant subset of explanatory variabl...
research
06/01/2016

Model selection consistency from the perspective of generalization ability and VC theory with an application to Lasso

Model selection is difficult to analyse yet theoretically and empiricall...
research
11/03/2022

Empirical Analysis of Model Selection for Heterogenous Causal Effect Estimation

We study the problem of model selection in causal inference, specificall...
research
06/10/2021

Problem-solving benefits of down-sampled lexicase selection

In genetic programming, an evolutionary method for producing computer pr...
research
10/12/2015

Toward a Better Understanding of Leaderboard

The leaderboard in machine learning competitions is a tool to show the p...
research
05/31/2022

The Environmental Discontinuity Hypothesis for Down-Sampled Lexicase Selection

Down-sampling training data has long been shown to improve the generaliz...

Please sign up or login with your details

Forgot password? Click here to reset