arrow
Return

Benchmarking saliency methods for chest X-ray interpretation

delete2022-10-10
delete78
delete
OA
AI
A
Adriel Saporta
X
Xiaotong Gui
A
Ashwin Agrawal
A
Anuj Pareek
S
Steven Q. H. Truong
C
Chanh D. Tr. Nguyen
V
Van-Doan Ngo
J
Jayne Seekins
F
Francis G. Blankenberg
A
Andrew Y. Ng
M
Matthew P. Lungren
P
Pranav Rajpurkar *
DOI:10.1038/s42256-022-00536-xdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Saliency methods, which produce heat maps that highlight the areas of the medical image that influence model prediction, are often presented to clinicians as an aid in diagnostic decision-making. However, rigorous investigation of the accuracy and reliability of these strategies is necessary before they are integrated into the clinical setting. In this work, we quantitatively evaluate seven saliency methods, including Grad-CAM, across multiple neural network architectures using two evaluation metrics. We establish the first human benchmark for chest X-ray segmentation in a multilabel classification set-up, and examine under what clinical conditions saliency maps might be more prone to failure in localizing important pathologies compared with a human expert benchmark. We find that (1) while Grad-CAM generally localized pathologies better than the other evaluated saliency methods, all seven performed significantly worse compared with the human benchmark, (2) the gap in localization performance between Grad-CAM and the human benchmark was largest for pathologies that were smaller in size and had shapes that were more complex, and (3) model confidence was positively correlated with Grad-CAM localization performance. Our work demonstrates that several important limitations of saliency methods must be addressed before we can rely on them for deep learning explainability in medical imaging. Saliency methods are used to localize areas of medical images that influence machine learning model predictions, but their accuracy and reliability require investigation. Saporta and colleagues evaluate seven saliency methods using different model architectures, and find that saliency maps perform worse than a human radiologist benchmark.
Keywords:
DEEP
MODEL

Journal

Nature Machine Intelligence cover
Nature Machine Intelligence
IF:
23.9
Papers:
1.3K
Citations:
1.5W

Organization

N
New York University
Scholars:
4.4W
Papers: 3.9W
Citations: 5.8W
S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
H
Harvard Medical School
Scholars:
6.5W
Papers: 4.8W
Citations: 91
V
VinUniversity
Scholars:
754
Papers: 419
Citations: 3
researcher View more organizations