Return
Improving crash data quality by identifying misclassified alcohol-involved crashes using NLP on narrative data
S
R
I
S
N
A
DOI:10.1016/j.jsr.2026.05.007.png)
Abstract
En 中文
Introduction: Road traffic crashes remain a leading cause of fatalities worldwide, underscoring the need for accurate data to guide prevention strategies and evidence-based policymaking. However, crash databases often suffer from misclassification, underreporting, and inconsistencies, particularly in alcohol-involved cases, which limits the reliability of safety analyses. Method: This study addresses this issue by identifying and quantifying Misclassified Alcohol-Involved Crashes (MAICs) using a Natural Language Processing (NLP) framework based on the BERT model. The framework analyzed 371,062 crash records from Iowa (2016-2022) and identified 3,895 misclassified alcohol-involved crashes (MAICs) out of 19,177 alcohol-involved cases predicted by the model, corresponding to an overall misclassification rate of 20.35% and a confidence interval of 18.86%-21.85%. To examine the factors contributing to these errors, a mixed-effects Probit Logit regression model was applied, incorporating behavioral, environmental, and roadway attributes. Results: Results indicated that fatal and nighttime crashes were less likely to be misclassified, whereas crashes involving older or younger drivers, heavy trucks, and vulnerable road users showed higher odds of misclassification. A Local Indicators of Spatial Association (LISA) analysis revealed significant county-level clusters of misclassifications, suggesting regional differences in enforcement and reporting practices.
Keywords:
Alcohol-involved crashes
Misclassification
Natural Language Processing
Crash data quality
Road safety analysis
Journal
IF:
4.4
Papers:
2.7K
Citations:
6.7K
