Return
LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews
L
B
M
DOI:10.1016/j.infsof.2026.108204.png)
Abstract
En 中文
• Use distilled good practices & recommendations for evaluating LLMs in SR screening. • Report confusion matrices enabling (meta-)analyses & alternative metric computation. • Prioritize lost evidence/recall in SR screening evaluations and Weighted MCC (WMCC). • Use cost–benefit analysis where lost evidence is a critical issue.
Keywords:
Large language models
LLM
Classification metrics
Class imbalance
Systematic reviews
Lost evidence
Cost-sensitive
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.3
Papers:
3.7K
Citations:
7.7K
