Importance of feature engineering and database selection in a machine learning model: A case study on carbon crystal structures

by   Franz M. Rohrhofer, et al.

Drive towards improved performance of machine learning models has led to the creation of complex features representing a database of condensed matter systems. The complex features, however, do not offer an intuitive explanation on which physical attributes do improve the performance. The effect of the database on the performance of the trained model is often neglected. In this work we seek to understand in depth the effect that the choice of features and the properties of the database have on a machine learning application. In our experiments, we consider the complex phase space of carbon as a test case, for which we use a set of simple, human understandable and cheaply computable features for the aim of predicting the total energy of the crystal structure. Our study shows that (i) the performance of the machine learning model varies depending on the set of features and the database, (ii) is not transferable to every structure in the phase space and (iii) depends on how well structures are represented in the database.


page 3

page 5

page 12

page 14

page 15

page 16

page 17


A Topological-Framework to Improve Analysis of Machine Learning Model Performance

As both machine learning models and the datasets on which they are evalu...

Computationally Efficient Feature Significance and Importance for Machine Learning Models

We develop a simple and computationally efficient significance test for ...

Classification of fetal compromise during labour: signal processing and feature engineering of the cardiotocograph

Cardiotocography (CTG) is the main tool used for fetal monitoring during...

Happy or Evil Laughter? Analysing a Database of Natural Audio Samples

We conducted a data collection on the basis of the Google AudioSet datab...

QuantifyML: How Good is my Machine Learning Model?

The efficacy of machine learning models is typically determined by compu...

Performance Prediction and Optimization of Solar Water Heater via a Knowledge-Based Machine Learning Method

Measuring the performance of solar energy and heat transfer systems requ...

Mill.jl and JsonGrinder.jl: automated differentiable feature extraction for learning from raw JSON data

Learning from raw data input, thus limiting the need for manual feature ...

Please sign up or login with your details

Forgot password? Click here to reset