arrow
Return

Automated chain-of-thought evaluation framework for large language model-generated emergency department documentation: a simulation-based study

delete2026-03-01
delete0
PRE
AI
D
Dasol Choi
J
Junhyuk Seo
W
Won Cul
K
Kim, Minha
H
Heo, Sejin
C
Chang, Hansol
K
Kim, Taerim *
DOI:10.15441/ceem.25.153delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Objective This study aimed to develop and validate MEDIVAL (Medical Documentation Validation), a progressive chain-of-thought (CoT) evaluation framework for automated assessment of large language model (LLM)-generated emergency department documentation, designed to align with expert clinical judgment in acute care settings. Methods We designed a three-tier evaluation framework incorporating persona-based, error-enhanced, and insight-integrated strategies. The framework was tested across four LLMs (GPT-4o, GPT-4.1, Claude-3.5, Claude-3.7) on 33 emergency department records reviewed by four expert emergency physicians. Each model applied the three CoT strategies across five criteria: appropriateness, accuracy, structure/format, conciseness, and clinical validity. Model outputs were compared with expert ratings using Spearman correlation coefficients. Differences were analyzed with the Friedman test and Wilcoxon signed rank test with Bonferroni correction. Reproducibility was assessed through intraclass correlation coefficient (ICC) analysis. Results All models demonstrated stronger alignment with expert ratings as CoT complexity increased, with Claude-3.7 (r=0.712, P<0.001) and GPT-4o (r=0.702, P<0.001) showing the highest correlations under the insight-integrated strategy. GPT-4.1 showed the greatest relative improvement (43.3% increase, r=0.457 to r=0.655, P<0.001). Significant overall differences were observed across strategies (chi(2)(2)=48.39, P<0.001), though the error-enhanced and insight-integrated approaches differed only modestly yet significantly (P=0.002). High reproducibility was confirmed (ICC >0.919), with Claude-3.5 achieving the most consistent results (ICC, 0.997-0.998). Conclusion MEDIVAL demonstrates that progressive CoT strategies systematically improve automated evaluation of emergency department documentation while maintaining excellent reproducibility. This framework offers a viable prescreening tool to reduce expert workload and support reliable artificial intelligence integration into emergency medicine workflows.
Keywords:
Artificial intelligence
Medical documentation
Emergency department
Large language models
Clinical evaluation

Journal

C
Clinical and Experimental Emergency Medicine
IF:
2.3
Papers:
25
Citations:
0

Organization

S
samsung
Scholars:
8.6K
Papers: 6.4K
Citations: 8
S
sungkyunkwan university (skku)
Scholars:
3.7W
Papers: 3.6W
Citations: 49
S
Samsung Medical Center
Scholars:
1.1W
Papers: 1.0W
Citations: 8.8K
Y
Yonsei University
Scholars:
4.8W
Papers: 4.6W
Citations: 5.2W
researcher View more organizations