1
Return

Large language models for systematic reviews were reported to perform well but rarely with verifiable safeguards: a cross-sectional study

delete2026-06-17
delete0
PRE
AI
H
Honghao Lai
B
Bernardo Sousa-Pinto
C
Christian Cao
D
David Moher
J
Janne Estill
J
Jiayi Liu
W
Weilong Zhao
Y
Yutong Wang
Z
Ziying Ye
B
Bo Tong
Z
Zhenhua Yang
X
Xufei Luo
B
Bingyi Wang
Y
Yimeng Li
P
Pan Bei
L
Lu Zhang
J
Jinhui Tian
Y
Yaolong Chen
N
Nannan Shi
葛龙 (Long Ge) *
DOI:10.1016/j.jclinepi.2026.112383delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• Mapped 229 studies (440 tasks) using transformer-based LLMs in systematic reviews • Assessed transparency, methodological quality, and how claims are communicated • Transparency was moderate (mean CHART score 0.52) but reproducibility gaps persist • Test-set locking unreported in 99.8% of tasks; leakage safeguards rarely verifiable • 60.6% of tasks were low risk of bias; practice-readiness claims frequent yet cautious

Journal

Journal of Clinical Epidemiology cover
Journal of Clinical Epidemiology
IF:
5.2
Papers:
8.4K
Citations:
4.3W

Organization

O
Ottawa Hospital Research Institute
Scholars:
6.5K
Papers: 5.2K
Citations: 1.3W
C
China Academy of Chinese Medical Sciences
Scholars:
8.6K
Papers: 4.4K
Citations: 1.1K
U
University of Porto
Scholars:
2.4K
Papers: 1.0K
Citations: 849
L
Lanzhou University
Scholars:
9.9K
Papers: 2.8K
Citations: 3.8W
U
university of toronto
Scholars:
14.5W
Papers: 11.9W
Citations: 165
H
hong kong baptist university
Scholars:
994
Papers: 586
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers