arrow
Return

Recommendations on compiling test datasets for evaluating artificial intelligence solutions in pathology

delete2022-12-01
delete33
delete
OA
AI
A
André Homeyer *
C
Christian Geißler
L
Lars Ole Schwen
F
Falk Zakrzewski
T
Theodore Evans
K
Klaus Strohmenger
M
Max Westphal
B
Buelow, Roman David
M
Michaela Kargl
A
Aray Karjauv
I
Isidre Munné-Bertran
C
Carl Orge Retzlaff
A
Adrià Romero-López
T
Tomasz Sołtysiński
M
Markus Plass
R
Rita Carvalho
S
Steinbach, Peter
Y
Yu-Chia Lan
N
Nassim Bouteldja
D
David Haber
M
Mateo Rojas-Carulla
A
Alireza Vafaei Sadr
M
Matthias Kraft
D
Daniel Krüger
R
Rutger Fick
T
Tobias Lang
P
Peter Boor
H
Heimo Müller
P
Peter Hufnagl
N
Norman Zerbe
DOI:10.1038/s41379-022-01147-ydelete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Artificial intelligence (AI) solutions that automatically extract information from digital histology images have shown great promise for improving pathological diagnosis. Prior to routine use, it is important to evaluate their predictive performance and obtain regulatory approval. This assessment requires appropriate test datasets. However, compiling such datasets is challenging and specific recommendations are missing. A committee of various stakeholders, including commercial AI developers, pathologists, and researchers, discussed key aspects and conducted extensive literature reviews on test datasets in pathology. Here, we summarize the results and derive general recommendations on compiling test datasets. We address several questions: Which and how many images are needed? How to deal with low-prevalence subsets? How can potential bias be detected? How should datasets be reported? What are the regulatory requirements in different countries? The recommendations are intended to help AI developers demonstrate the utility of their products and to help pathologists and regulatory agencies verify reported performance measures. Further research is needed to formulate criteria for sufficiently representative test datasets so that AI solutions can operate with less user intervention and better support diagnostic workflows in the future.
Keywords:
INTEROBSERVER VARIABILITY
BREAST-CANCER
VALIDATION
PREDICTION
CONSENSUS
IMAGES
MODEL
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Modern Pathology cover
Modern Pathology
IF:
5.5
Papers:
5.3K
Citations:
1.8W

Organization

O
olympus corporation
Scholars:
190
Papers: 123
Citations: 2
R
RWTH Aachen University
Scholars:
3.5W
Papers: 2.6W
Citations: 3.6W
B
Berlin Institute of Health
Scholars:
3.9W
Papers: 3.0W
Citations: 6.6K
M
Medical University of Graz
Scholars:
1.4W
Papers: 9.9K
Citations: 1.2W
F
Free University of Berlin
Scholars:
3.8W
Papers: 3.2W
Citations: 51
T
Technical University of Berlin
Scholars:
1.3W
Papers: 1.1W
Citations: 18
R
RWTH Aachen University Hospital
Scholars:
4.8K
Papers: 3.7K
Citations: 5
H
Helmholtz Association
Scholars:
13.2W
Papers: 10.7W
Citations: 145
H
helmholtz-zentrum dresden-rossendorf (hzdr)
Scholars:
3.5K
Papers: 2.3K
Citations: 2
H
Humboldt University of Berlin
Scholars:
3.2W
Papers: 2.7W
Citations: 47
researcher View more organizations