Return
Large language models for systematic reviews were reported to perform well but rarely with verifiable safeguards: a cross-sectional study
H
B
C
D
J
J
W
Y
Z
B
Z
X
B
Y
P
L
J
Y
N
葛
DOI:10.1016/j.jclinepi.2026.112383.png)
Abstract
En 中文
• Mapped 229 studies (440 tasks) using transformer-based LLMs in systematic reviews • Assessed transparency, methodological quality, and how claims are communicated • Transparency was moderate (mean CHART score 0.52) but reproducibility gaps persist • Test-set locking unreported in 99.8% of tasks; leakage safeguards rarely verifiable • 60.6% of tasks were low risk of bias; practice-readiness claims frequent yet cautious
Journal
IF:
5.2
Papers:
8.4K
Citations:
4.3W
