1
Return

Measuring the quality of AI-generated clinical notes: A systematic review and experimental benchmark of evaluation methods

delete2026-04-07
delete0
delete
OA
AI
A
Alexandra Dahlberg *
T
Tiila Käenniemi
T
Tiia Winther-Jensen
O
Olli Tapiola
R
Rami Luisto
T
Tuukka Puranen
M
Max Gordon
E
Enni Sanmark
V
Ville Vartiainen
DOI:10.1016/j.artmed.2026.103421delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• Systematic review shows clinical note evaluation relies mainly on ROUGE and BLEU • Lexical overlap metrics mis-rank meaning-preserving paraphrases in clinical notes • Controlled perturbation framework benchmarks evaluation metrics for clinical text • Semantic and LLM-based evaluators better detect clinically relevant content changes • Evaluation outcomes vary by language and model capacity, limiting generalisability
Keywords:
Artificial intelligence in medicine
Clinical natural language processing
Large LANGUAGE models
Clinical documentation
Text quality evaluation
Semantic similarity
Automated evaluation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Artificial Intelligence in Medicine cover
Artificial Intelligence in Medicine
IF:
6.2
Papers:
2.5K
Citations:
7.8K

Organization

U
university of jyvaskyla
Scholars:
6.3K
Papers: 6.8K
Citations: 12
U
University of Helsinki
Scholars:
4.2K
Papers: 1.7K
Citations: 5.1W
K
Karolinska Institutet
Scholars:
5.8W
Papers: 4.8W
Citations: 7.1W
Cited Papers

Cited Papers

Citing Papers

Citing Papers