arrow
Return

Explainable AI for Coronary Artery Disease Stratification Using Routine Clinical Data

delete2025-11-03
delete0
PRE
AI
N
Nurdaulet Tasmurzayev
Б
Баглан Иманбек
A
Assiya Boltaboyeva *
G
Gulmira Dikhanbayeva
S
Sarsenbek Zhussupbekov
Q
Qarlygash Saparbayeva
G
Gulshat Amirkhanova
DOI:10.3390/a18110693delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Background: Coronary artery disease (CAD) remains a leading cause of morbidity and mortality. Early diagnosis reduces adverse outcomes and alleviates the burden on healthcare, yet conventional approaches are often invasive, costly, and not always available. In this context, machine learning offers promising solutions. Objective: The objective of this study is to evaluate the feasibility of reliably predicting both the presence and the severity of CAD. The analysis is based on a harmonized, multi-center UCI dataset that includes cohorts from Cleveland, Hungary, Switzerland, and Long Beach. The work aims to assess the accuracy and practical utility of models built exclusively on routine tabular clinical and demographic data, without relying on imaging. These models are designed to improve risk stratification and guide patient routing. Methods and Results: The study is based on a uniform and standardized data processing pipeline. This pipeline includes handling missing values, feature encoding, scaling, an 80/20 train-test split and applying the SMOTE method exclusively to the training set to prevent information leakage. Within this pipeline, a standardized comparison of a wide range of models (including gradient boosting, tree-based ensembles, support vector methods, etc.) was conducted with hyperparameter tuning via GridSearchCV. The best results were demonstrated by the CatBoost model: accuracy-0.8278, recall-0.8407, and F1-score-0.8436. Conclusions: A key distinction of this work is the comprehensive evaluation of the models' practical suitability. Beyond standard metrics, the analysis of calibration curves confirmed the reliability of the probabilistic predictions. Patient-level interpretability using SHAP showed that the model relies on clinically significant predictors, including ST-segment depression. Calibrated and explainable models based on readily available data are positioned as a practical tool for scalable risk stratification and decision support, especially in resource-constrained settings.
Keywords:
coronary artery disease
machine learning
CatBoost
CVD
risk prediction
ROC-AUC

Journal

Algorithms cover
Algorithms
IF:
2.1
Papers:
631
Citations:
5.4K

Organization

A
al-farabi kazakh national university
Scholars:
563
Papers: 214
Citations: 0