ISSN: 1304-7191 | E-ISSN: 1304-7205
Diabetes onset prediction using random forest: A machine learning approach with the Pima Indians diabetes dataset
1Department of Information Technology, Eastern Visayas State University, Leyte, 6500, Philippines
2Department of Information Technology, Visayas State University, Leyte, 6521, Philippines
Sigma J Eng Nat Sci 2026; 44(3): 1816-1825 DOI: 10.14744/sigma.2026.2045
Full Text PDF

Abstract

Diabetes, a chronic metabolic disorder, has affected millions of people worldwide, thus it is important to develop accurate predictive models for early intervention and improved patient prognosis. This paper aims to introduce a predictive model for diabetes onset using the Pima Indians Diabetes Dataset and the random forest algorithm, a machine learning approach that handles complex, nonlinear data well. The model was built with the goal of optimizing accu-racy, reliability and interpretability for clinical applications. Data normalization and class im-balance management with the synthetic minority over-sampling Technique were among the data pre-processing steps that resulted in a substantial improvement in model performance. The random forest model exhibited good performance in predicting diabetes with an accuracy of 80%, precision of 77% and recall of 68%. The feature importance analysis revealed that glu-cose levels, body mass index and age are the most important predictors. To improve the model interpretability and clinical utility, the predictions were interpreted using the Shapley Additive Explanations, thereby allowing the healthcare professionals to trust and utilize the model for risk assessment. The results underscore the potential of machine learning, particularly the random forest algorithm in aiding early diagnosis of diabetes and assisting healthcare provid-ers to identify high-risk individuals for timely intervention.