1
Return

LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews

delete2026-06-08
delete0
delete
OA
AI
L
Lech Madeyski *
B
Barbara Kitchenham
M
Martin Shepperd
DOI:10.1016/j.infsof.2026.108204delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• Use distilled good practices & recommendations for evaluating LLMs in SR screening. • Report confusion matrices enabling (meta-)analyses & alternative metric computation. • Prioritize lost evidence/recall in SR screening evaluations and Weighted MCC (WMCC). • Use cost–benefit analysis where lost evidence is a critical issue.
Keywords:
Large language models
LLM
Classification metrics
Class imbalance
Systematic reviews
Lost evidence
Cost-sensitive
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Information and Software Technology cover
Information and Software Technology
IF:
4.3
Papers:
3.7K
Citations:
7.7K

Organization

B
brunel university
Scholars:
5.8K
Papers: 7.0K
Citations: 9
K
keele university
Scholars:
464
Papers: 259
Citations: 0
W
wrocław university of science and technology
Scholars:
216
Papers: 106
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers