Predict Diabetes Using Voting Classifier and Hyper Tuning Technique
Today, diabetes is one of the most common chronic diseases in the world due to the people’s sedentary lifestyle which led to many health issues like heart attack, kidney frailer and blindness. Additionally, most of the people are unrealizable about the early-stage diabetes symptoms to prevent it. Th...
Saved in:
Main Authors: | , |
---|---|
Format: | Article |
Language: | English |
Published: |
Sulaimani Polytechnic University
2023-01-01
|
Series: | Kurdistan Journal of Applied Research |
Subjects: | |
Online Access: | https://kjar.spu.edu.iq/index.php/kjar/article/view/821 |
Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
_version_ | 1823861374601658368 |
---|---|
author | Chra Ali Kamal Manal Ali Atiyah |
author_facet | Chra Ali Kamal Manal Ali Atiyah |
author_sort | Chra Ali Kamal |
collection | DOAJ |
description | Today, diabetes is one of the most common chronic diseases in the world due to the people’s sedentary lifestyle which led to many health issues like heart attack, kidney frailer and blindness. Additionally, most of the people are unrealizable about the early-stage diabetes symptoms to prevent it. The above reasons were encouraging to develop a diabetes prediction system using machine learning techniques. The Pima Indian Diabetes Dataset (PIDD) was utilized for this framework as it is common and appropriate dataset in .CSV format. While there were not any duplicate or null values, however, some zero values were replaced, four outlier records were removed and data standardization were performed in the dataset. In addition, this project methodology divided into two phases of model selection. In the first phase, two different hyper parameter techniques (Randomized Search and TPOT(autoML)) were used to increase the accuracy level for each algorithm. Then six different algorithms (Logistic Regression, Decision Tree, Random Forest, K-nearest neighbor, Support Vector Machine and Naïve Bayes) were applied. In the second phase, the four best performed algorithms (with best estimated parameters for each of them) were chosen and used as an input for the voting classifier, because it applies to find the best algorithm between a group of multiple options. The result was satisfying, and Random Forest was achieved 98.69% in second stage, while its accuracy level was 81.04% in the previous one and it utilized to predict diabetes via a simple graphic user interface.
|
format | Article |
id | doaj-art-e07e73b3e4f545a6a26c29f8bf78da5b |
institution | Kabale University |
issn | 2411-7684 2411-7706 |
language | English |
publishDate | 2023-01-01 |
publisher | Sulaimani Polytechnic University |
record_format | Article |
series | Kurdistan Journal of Applied Research |
spelling | doaj-art-e07e73b3e4f545a6a26c29f8bf78da5b2025-02-09T20:59:38ZengSulaimani Polytechnic UniversityKurdistan Journal of Applied Research2411-76842411-77062023-01-017210.24017/Science.2022.2.10Predict Diabetes Using Voting Classifier and Hyper Tuning TechniqueChra Ali Kamal0Manal Ali Atiyah1Database Department, Computer Science Institute, Sulaimani, Polytechnic University Sulaimani,IraqDatabase Department, Computer Science Institute, Sulaimani Polytechnic University, Sulaimani, IraqToday, diabetes is one of the most common chronic diseases in the world due to the people’s sedentary lifestyle which led to many health issues like heart attack, kidney frailer and blindness. Additionally, most of the people are unrealizable about the early-stage diabetes symptoms to prevent it. The above reasons were encouraging to develop a diabetes prediction system using machine learning techniques. The Pima Indian Diabetes Dataset (PIDD) was utilized for this framework as it is common and appropriate dataset in .CSV format. While there were not any duplicate or null values, however, some zero values were replaced, four outlier records were removed and data standardization were performed in the dataset. In addition, this project methodology divided into two phases of model selection. In the first phase, two different hyper parameter techniques (Randomized Search and TPOT(autoML)) were used to increase the accuracy level for each algorithm. Then six different algorithms (Logistic Regression, Decision Tree, Random Forest, K-nearest neighbor, Support Vector Machine and Naïve Bayes) were applied. In the second phase, the four best performed algorithms (with best estimated parameters for each of them) were chosen and used as an input for the voting classifier, because it applies to find the best algorithm between a group of multiple options. The result was satisfying, and Random Forest was achieved 98.69% in second stage, while its accuracy level was 81.04% in the previous one and it utilized to predict diabetes via a simple graphic user interface. https://kjar.spu.edu.iq/index.php/kjar/article/view/821Diabetic PredictionGraphical User InterfaceHyper TuningMachine Learning AlgorithmVoting Classifier |
spellingShingle | Chra Ali Kamal Manal Ali Atiyah Predict Diabetes Using Voting Classifier and Hyper Tuning Technique Kurdistan Journal of Applied Research Diabetic Prediction Graphical User Interface Hyper Tuning Machine Learning Algorithm Voting Classifier |
title | Predict Diabetes Using Voting Classifier and Hyper Tuning Technique |
title_full | Predict Diabetes Using Voting Classifier and Hyper Tuning Technique |
title_fullStr | Predict Diabetes Using Voting Classifier and Hyper Tuning Technique |
title_full_unstemmed | Predict Diabetes Using Voting Classifier and Hyper Tuning Technique |
title_short | Predict Diabetes Using Voting Classifier and Hyper Tuning Technique |
title_sort | predict diabetes using voting classifier and hyper tuning technique |
topic | Diabetic Prediction Graphical User Interface Hyper Tuning Machine Learning Algorithm Voting Classifier |
url | https://kjar.spu.edu.iq/index.php/kjar/article/view/821 |
work_keys_str_mv | AT chraalikamal predictdiabetesusingvotingclassifierandhypertuningtechnique AT manalaliatiyah predictdiabetesusingvotingclassifierandhypertuningtechnique |