The Use of Hellinger Distance Undersampling Model to Improve the Classification of Disease Class in Imbalanced Medical Datasets

Imbalanced class distribution in the medical dataset is a challenging task that hinders classifying disease correctly. It emerges when the number of healthy class instances being much larger than the disease class instances. To solve this problem, we proposed undersampling the healthy class instance...

Full description

Saved in:
Bibliographic Details
Main Authors: Zina Z. R. Al-Shamaa, Sefer Kurnaz, Adil Deniz Duru, Nadia Peppa, Alex H. Mirnezami, Zaed Z. R. Hamady
Format: Article
Language:English
Published: Wiley 2020-01-01
Series:Applied Bionics and Biomechanics
Online Access:http://dx.doi.org/10.1155/2020/8824625
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:Imbalanced class distribution in the medical dataset is a challenging task that hinders classifying disease correctly. It emerges when the number of healthy class instances being much larger than the disease class instances. To solve this problem, we proposed undersampling the healthy class instances to improve disease class classification. This model is named Hellinger Distance Undersampling (HDUS). It employs the Hellinger Distance to measure the resemblance between majority class instance and its neighbouring minority class instances to separate classes effectively and boost the discrimination power for each class. An extensive experiment has been conducted on four imbalanced medical datasets using three classifiers to compare HDUS with a baseline model and three state-of-the-art undersampling models. The outcomes display that HDUS can perform better than other models in terms of sensitivity, F1 measure, and balanced accuracy.
ISSN:1176-2322
1754-2103