EXPLAINABLE MACHINE LEARNING FOR EARLY PREDICTION OF CARDIOVASCULAR DISEASE USING CLINICAL AND DEMOGRAPHIC DATA
Keywords:
Machine learning; Cardiovascular disease; Explainable AI; XGBoost; SHAP; Healthcare analytics.Abstract
Given that cardiovascular diseases are the leading cause of morbidity and mortality in the world, it is important to predict the disease in the early stages from a medical perspective. These machine learning methods can identify complex relationships within the clinical and demographic variables and assist in the early prediction of the disease. However, problems arise regarding the interpretability of many prediction models and their implementation into the practice of medicine. Therefore, the current project aimed at developing an interpretable machine learning algorithm for the early prediction of the disease and its evaluation in terms of predictivity, interpretability, and computational efficiency. In this case, the dataset included 1,024 patients' records with 13 clinical and demographic features. The process of missing value imputation, normalization, and feature selection took place in the process of data preprocessing. The Logistic Regression, Random Forest, Support Vector Machine, and XGBoost algorithms were developed and examined using 80% of training and 20% of testing datasets. Accuracy, Precision, Recall, F1-Score, and Area under ROC Curve (AUC) were taken as the criteria for evaluating the models. Explainability was done using the SHAP method, which helped identify the most important predictors. As far as the prediction results go, XGBoost algorithm gave the highest scores on accuracy (91.2%), precision (90.5%), recall (89.8%) and F1-score (90.1%) compared to other algorithms where accuracy varied from 84.7% to 88.9%. XGBoost algorithm gave an AUC score of 0.94. According to SHAP analysis, the important predictors for predicting cardiovascular diseases are age (18.6%), cholesterol (16.9%), max heart rate (14.7%), and resting blood pressure (13.5%). The proposed approach outperformed the baseline ensemble algorithm in terms of prediction time by 32%.


