Detecting Chronic Kidney Disease(CKD) at the Initial Stage: A Novel Hybrid Feature-selection Method and Robust Data Preparation Pipeline for Different ML Techniques

Chronic Kidney Disease (CKD) has infected almost 800 million people around the world. Around 1.7 million people die each year because of it. Detecting CKD in the initial stage is essential for saving millions of lives. Many researchers have applied distinct Machine Learning (ML) methods to detect CKD at an early stage, but detailed studies are still missing. We present a structured and thorough method for dealing with the complexities of medical data with optimal performance. Besides, this study will assist researchers in producing clear ideas on the medical data preparation pipeline. In this paper, we applied KNN Imputation to impute missing values, Local Outlier Factor to remove outliers, SMOTE to handle data imbalance, K-stratified K-fold Cross-validation to validate the ML models, and a novel hybrid feature selection method to remove redundant features. Applied algorithms in this study are Support Vector Machine, Gaussian Naive Bayes, Decision Tree, Random Forest, Logistic Regression, K-Nearest Neighbor, Gradient Boosting, Adaptive Boosting, and Extreme Gradient Boosting. Finally, the Random Forest can detect CKD with 100

READ FULL TEXT

page 5

page 7

research
06/18/2021

Performance Evaluation of Classification Models for Household Income, Consumption and Expenditure Data Set

Food security is more prominent on the policy agenda today than it has b...
research
02/28/2022

An empirical comparison of machine learning models for student's mental health illness assessment

Student's mental health problems have been explored previously in higher...
research
02/28/2021

Machine learning for detection of stenoses and aneurysms: application in a physiologically realistic virtual patient database

This study presents an application of machine learning (ML) methods for ...
research
06/09/2022

Human Activity Recognition from Knee Angle Using Machine Learning Techniques

Human Activity Recognition (HAR) is a crucial technology for many applic...
research
07/02/2020

A Machine Learning Pipeline Stage for Adaptive Frequency Adjustment

A machine learning (ML) design framework is proposed for adaptively adju...

Please sign up or login with your details

Forgot password? Click here to reset