Return
Traffic Accident Prediction Considering Imbalanced Data: An Interpretable Machine Learning Model
DOI:10.1007/s12555-025-0032-7.png)
Abstract
En 中文
Traffic accident prediction is vital for enhancing road safety and mitigating accident risks. However, the inherent imbalance in accident data poses significant challenges to model performance, while existing models often neglect the temporal heterogeneity of holidays and peak hours and suffer from limited interpretability. This study presents a comprehensive traffic accident prediction framework to address these gaps. First, a novel data generation approach, incorporating gradient penalty terms into the Wasserstein generative adversarial network, is proposed to alleviate data imbalance and generate high-quality synthetic data. Second, time classification features, capturing the effects of time of day, holidays, and peak hours, are integrated into an enhanced extreme gradient boosting model for accident prediction. Lastly, the Shapley additive explanations method is applied to interpret the model results and uncover the nonlinear relationships between key factors and accident occurrence. Using accident and traffic data from the Xi'an ring expressway, the framework demonstrates superior performance over traditional methods in accuracy and reliability. Notably, average speed emerges as the most critical factor influencing accident risk, with a heightened risk observed when the mixing degree falls below 0.22 during peak hours or exceeds 0.7 universally. These findings provide actionable insights for proactive traffic safety management, enabling precise risk assessments and informed decision-making.
Keywords:
Data imbalance
interpretability
machine learning
time classification features
traffic accident prediction
Journal
IF:
2.9
Papers:
216
Citations:
6.5K

