2Department of Computer Engineering, K.K.Wagh Institute of Engineering Education and Research, Savitribai Phule Pune University, 422003, India
Abstract
The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF). This does not require siloing such methods as SMOTE or NearMiss, but it does change the way to create or make samples, as well as the way to achieve retention, depending on group size. We evaluate the performance of each strategy by considering rare cases, using Extra Trees Classifier. We ran tests on the following data collections: CDC Diabetes, Breast Cancer Wisconsin (Diagnostic), and Credit Card Fraud Detection, KDD Cup 1999 Intrusion Detection, PIMA, and BRFSS. In the performance trend, ASCPF demonstrated the highest accuracy, no matter compared to SMOTE or NearMiss. Take CDC Diabetes. Here, ASCPF got to 94.12% in accuracy, pulled 87.99% sensitivity, hit 94.10% ROC-AUC - way ahead of SMOTE’s 77.65% accuracy and NearMiss’s 90.47%. It was important to keep the original proportions of groups to avoid to create misbalanced distribution and without changing the results it helped to achieve adding or removing the frequencies for cases of very few occurrences. One distinctive feature is that it is a combination of creating smart fake samples and ratio control – a novel approach which makes ASCPF unique. The combination results in more uniform performance on classes with skewed distribution, especially in important health decisions.
