Return
Measuring the quality of AI-generated clinical notes: A systematic review and experimental benchmark of evaluation methods
A
T
T
O
R
T
M
E
V
DOI:10.1016/j.artmed.2026.103421.png)
Abstract
En 中文
• Systematic review shows clinical note evaluation relies mainly on ROUGE and BLEU • Lexical overlap metrics mis-rank meaning-preserving paraphrases in clinical notes • Controlled perturbation framework benchmarks evaluation metrics for clinical text • Semantic and LLM-based evaluators better detect clinically relevant content changes • Evaluation outcomes vary by language and model capacity, limiting generalisability
Keywords:
Artificial intelligence in medicine
Clinical natural language processing
Large LANGUAGE models
Clinical documentation
Text quality evaluation
Semantic similarity
Automated evaluation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
6.2
Papers:
2.5K
Citations:
7.8K
