arrow
返回

Benchmarking saliency methods for chest X-ray interpretation

delete2022-10-10
delete78
delete
OA
AI
A
Adriel Saporta
X
Xiaotong Gui
A
Ashwin Agrawal
A
Anuj Pareek
S
Steven Q. H. Truong
C
Chanh D. Tr. Nguyen
V
Van-Doan Ngo
J
Jayne Seekins
F
Francis G. Blankenberg
A
Andrew Y. Ng
M
Matthew P. Lungren
P
Pranav Rajpurkar *
DOI:10.1038/s42256-022-00536-xdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Saliency methods, which produce heat maps that highlight the areas of the medical image that influence model prediction, are often presented to clinicians as an aid in diagnostic decision-making. However, rigorous investigation of the accuracy and reliability of these strategies is necessary before they are integrated into the clinical setting. In this work, we quantitatively evaluate seven saliency methods, including Grad-CAM, across multiple neural network architectures using two evaluation metrics. We establish the first human benchmark for chest X-ray segmentation in a multilabel classification set-up, and examine under what clinical conditions saliency maps might be more prone to failure in localizing important pathologies compared with a human expert benchmark. We find that (1) while Grad-CAM generally localized pathologies better than the other evaluated saliency methods, all seven performed significantly worse compared with the human benchmark, (2) the gap in localization performance between Grad-CAM and the human benchmark was largest for pathologies that were smaller in size and had shapes that were more complex, and (3) model confidence was positively correlated with Grad-CAM localization performance. Our work demonstrates that several important limitations of saliency methods must be addressed before we can rely on them for deep learning explainability in medical imaging. Saliency methods are used to localize areas of medical images that influence machine learning model predictions, but their accuracy and reliability require investigation. Saporta and colleagues evaluate seven saliency methods using different model architectures, and find that saliency maps perform worse than a human radiologist benchmark.
Keyword:
DEEP
MODEL

期刊

Nature Machine Intelligence 封面图
Nature Machine Intelligence
IF:
23.9
论文数:
1.3K
被引数:
1.5W

机构

N
New York University
学者数:
4.4W
论文数: 3.9W
被引数: 5.8W
S
Stanford University
学者数:
9.6W
论文数: 8.2W
被引数: 17.0W
H
Harvard Medical School
学者数:
6.5W
论文数: 4.8W
被引数: 91
V
VinUniversity
学者数:
770
论文数: 430
被引数: 3
学者 查看更多机构
引用论文

引用论文

Lithocholic acid down-regulation of NF-κB activity through vitamin D receptor in colonic cancer cells
err2008-07-01
err0
errOAAI
errJun Sun; Reba Mustafi; Sonia Cerda; Anusara Chumsangsri; Yinglin Rick Xia; Yan Chun Li; Marc Bissonnette
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Evaluation of bioactive compounds and antibacterial activity of Pulicaria jaubertii extract obtained by supercritical and conventional methods
err2020-09-14
err0
PREAI
errQais Ali Al-Maqtari; Amer Ali Mahdi; Waleed Al‑Ansi; Jalaleldeen Khaleel Mohammed; Minping Wei; Weirong Yao
err分享
err收藏
Top-Down Neural Attention by Excitation Backprop自上而下的神经注意力通过激励背prop
err2017-12-23
err413
PREAI
errZhang, Jianming; Bargal, Sarah Adel; Lin, Zhe; Brandt, Jonathan; Shen, Xiaohui; Sclaroff, Stan
err分享
err收藏
AppendiXNet: Deep Learning for Diagnosis of Appendicitis from A Small Dataset of CT Exams Using Video Pretraining附录网: 使用视频预训练从ct检查的小数据集进行阑尾炎诊断的深度学习
err2020-03-03
err74
errOAAI
errRajpurkar, Pranav; Park, Allison; Irvin, Jeremy; Chute, Chris; Bereket, Michael; Mastrodicasa, Domenico; Langlotz, Curtis P.; Lungren, Matthew P.; Ng, Andrew Y.; Patel, Bhavik N.
err分享
err收藏
Subspheroids in the lithic assemblage of Barranco León (Spain): Recognizing the late Oldowan in Europe
err2020-01-30
err0
errOAAI
errStefania Titton; Deborah Barsky; Amèlia Bargalló; Alexia Serrano-Ramos; Josep Maria Vergès; Isidro Toro-Moyano; Robert Sala-Ramos; José García Solano; Juan Manuel Jimenez Arenas
err分享
err收藏
Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practicing radiologists胸部x光片诊断的深度学习: CheXNeXt算法与实践放射科医生的回顾性比较
err2018-11-20
err728
errOAAI
errRajpurkar, Pranav; Irvin, Jeremy; Ball, Robyn L.; Zhu, Kaylie; Yang, Brandon; Mehta, Hershel; Duan, Tony; Ding, Daisy; Bagul, Aarti; Langlotz, Curtis P.; Patel, Bhavik N.; Yeom, Kristen W.; Shpanskaya, Katie; Blankenberg, Francis G.; Seekins, Jayne; Amrhein, Timothy J.; Mong, David A.; Halabi, Safwan S.; Zucker, Evan J.; Ng, Andrew Y.; Lungren, Matthew P.
err分享
err收藏
学者 查看更多内容