1
Return

Predicting heart failure using MIMIC-IV and MIMIC-IV-ED: a comparative study of machine learning and deep learning models

delete2026-04-06
delete0
PRE
AI
T
Teoh, Jing Ru
H
Hasikin, Khairunnisa
N
Ng, Wei Lin
L
Lee, Kee Wei
W
Wu, Xiang
L
Lai, KW
DOI:10.7717/peerj-cs.3743delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Heart failure (HF) is a leading cause of morbidity and mortality worldwide, yet early identification remains a challenge in clinical practice. This study presents a comprehensive comparison of machine learning (ML) and deep learning (DL) models for HF prediction using large-scale, International Classification of Diseases (ICD)-coded cohorts derived from the MIMIC-IV and MIMIC-IV-ED databases. A total of 17,892 adult patients were included, comprising 9,373 HF cases and 8,519 non-HF controls. Structured clinical variables were extracted from electronic health records (EHR), while unstructured admission notes were encoded using BioClinicalBERT and integrated via dimensionality reduction. Rigorous preprocessing was applied, including outlier handling, feature selection guided by clinical guidelines, data imputation, and Z-score standardization. We evaluated multiple ML models (Random Forest (RF), Logistic Regression (LR), Decision Tree (DT), Na & iuml;ve Bayes (NB), adaptive boosting) and DL models (Multilayer Perceptron (MLP), Convolutional Neural Network (CNN), Dense Neural Network (DNN)). Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic (AUROC) with 95% confidence intervals (CI) estimated via bootstrap resampling. Among all models, Random Forest (RF) achieved the best overall performance (accuracy = 0.8994, AUROC = 0.9659), consistently outperforming both traditional ML and DL approaches. An ablation study demonstrated that incorporating bidirectional encoder representations from transformers (BERT)-derived admission note embeddings substantially improved RF performance but yielded more modest gains for neural models. A sensitivity analysis including patients with other heart-related conditions in the control group confirmed the robustness of the proposed framework under more clinically realistic settings. Overall, the findings of this study highlight the strong utility of ensemble-based ML models for structured EHR data and provide empirical insights into the model-dependent value of unstructured clinical text for HF risk stratification. Future work should validate these models externally and explore more advanced multimodal fusion strategies for clinical deployment.
Keywords:
Heart failure prediction
Electronic health records (EHR)
Machine learning
Deep learning
Comparative analysis
Multimodal clinical data

Journal

PeerJ Computer Science cover
PeerJ Computer Science
IF:
2.5
Papers:
3.3K
Citations:
6.9K

Organization

U
universiti malaya
Scholars:
3.5K
Papers: 1.5K
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers