arrow
Return

Comparative Forecasting and Misclassification Analysis Using Health Survey Data

delete2026-04-20
delete0
delete
OA
AI
E
Ermioni Traka
G
George Papageorgiou
G
Georgios Mantzavinis
C
Christos Tjortjis *
DOI:10.3390/ai7040148delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Background: Accurate mortality prediction remains a major challenge in public health due to the complex interactions among demographic, socioeconomic, behavioral, and medical factors. This problem is particularly relevant for identifying high-risk groups and improving preventive healthcare strategies. While existing studies demonstrate strong predictive performance, they mainly rely on clinically structured data and focus on model performance. Challenges such as misclassification and atypical cases remain less explored. Methods: Using the Integrated Public Use Microdata Series National Health Interview Survey (IPUMS-NHIS) 2010 and 2015 datasets (193,765 records, 104 features), this study investigates mortality prediction through comparative Machine Learning. Data preprocessing included feature engineering, categorical encoding, and removal of missing entries. Class imbalance was addressed using SMOTE and SMOTE-ENN resampling, followed by hyperparameter tuning. Three models—Logistic Regression, Random Forest, and XGBoost—were trained to classify mortality, with recall prioritized to ensure accurate identification of deceased cases. Results: Results showed that XGBoost achieved the best performance (Recall = 69%, F1 = 0.39, AUC = 0.92), outperforming other models in balancing sensitivity and specificity. Feature importance and permutation analyses highlighted age, employment status, self-reported health, and lifestyle indicators as key predictors. Misclassification analysis combined with Isolation Forest revealed atypical profiles not captured by standard models. Conclusions: The findings underscore XGBoost’s effectiveness and demonstrate the value of integrating anomaly detection with classification to improve mortality prediction and inform public health planning.
Keywords:
anomaly detection
class imbalance
logistic regression
machine learning
mortality prediction
public health forecasting
random forest
XGBoost
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

A
AI
IF:
5
Papers:
991
Citations:
941

Organization

I
international hellenic university
Scholars:
165
Papers: 102
Citations: 0
U
university of ioannina
Scholars:
588
Papers: 294
Citations: 0