Return
Predicting heart failure using MIMIC-IV and MIMIC-IV-ED: a comparative study of machine learning and deep learning models
T
H
N
L
W
L
DOI:10.7717/peerj-cs.3743.png)
Abstract
En 中文
Heart failure (HF) is a leading cause of morbidity and mortality worldwide, yet early identification remains a challenge in clinical practice. This study presents a comprehensive comparison of machine learning (ML) and deep learning (DL) models for HF prediction using large-scale, International Classification of Diseases (ICD)-coded cohorts derived from the MIMIC-IV and MIMIC-IV-ED databases. A total of 17,892 adult patients were included, comprising 9,373 HF cases and 8,519 non-HF controls. Structured clinical variables were extracted from electronic health records (EHR), while unstructured admission notes were encoded using BioClinicalBERT and integrated via dimensionality reduction. Rigorous preprocessing was applied, including outlier handling, feature selection guided by clinical guidelines, data imputation, and Z-score standardization. We evaluated multiple ML models (Random Forest (RF), Logistic Regression (LR), Decision Tree (DT), Na & iuml;ve Bayes (NB), adaptive boosting) and DL models (Multilayer Perceptron (MLP), Convolutional Neural Network (CNN), Dense Neural Network (DNN)). Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic (AUROC) with 95% confidence intervals (CI) estimated via bootstrap resampling. Among all models, Random Forest (RF) achieved the best overall performance (accuracy = 0.8994, AUROC = 0.9659), consistently outperforming both traditional ML and DL approaches. An ablation study demonstrated that incorporating bidirectional encoder representations from transformers (BERT)-derived admission note embeddings substantially improved RF performance but yielded more modest gains for neural models. A sensitivity analysis including patients with other heart-related conditions in the control group confirmed the robustness of the proposed framework under more clinically realistic settings. Overall, the findings of this study highlight the strong utility of ensemble-based ML models for structured EHR data and provide empirical insights into the model-dependent value of unstructured clinical text for HF risk stratification. Future work should validate these models externally and explore more advanced multimodal fusion strategies for clinical deployment.
Keywords:
Heart failure prediction
Electronic health records (EHR)
Machine learning
Deep learning
Comparative analysis
Multimodal clinical data
Journal
IF:
2.5
Papers:
3.3K
Citations:
6.9K
