返回
Benchmarking saliency methods for chest X-ray interpretation
DOI:10.1038/s42256-022-00536-x.png)
摘要
En 中文
Saliency methods, which produce heat maps that highlight the areas of the medical image that influence model prediction, are often presented to clinicians as an aid in diagnostic decision-making. However, rigorous investigation of the accuracy and reliability of these strategies is necessary before they are integrated into the clinical setting. In this work, we quantitatively evaluate seven saliency methods, including Grad-CAM, across multiple neural network architectures using two evaluation metrics. We establish the first human benchmark for chest X-ray segmentation in a multilabel classification set-up, and examine under what clinical conditions saliency maps might be more prone to failure in localizing important pathologies compared with a human expert benchmark. We find that (1) while Grad-CAM generally localized pathologies better than the other evaluated saliency methods, all seven performed significantly worse compared with the human benchmark, (2) the gap in localization performance between Grad-CAM and the human benchmark was largest for pathologies that were smaller in size and had shapes that were more complex, and (3) model confidence was positively correlated with Grad-CAM localization performance. Our work demonstrates that several important limitations of saliency methods must be addressed before we can rely on them for deep learning explainability in medical imaging. Saliency methods are used to localize areas of medical images that influence machine learning model predictions, but their accuracy and reliability require investigation. Saporta and colleagues evaluate seven saliency methods using different model architectures, and find that saliency maps perform worse than a human radiologist benchmark.
Keyword:
DEEP
MODEL
期刊
IF:
23.9
论文数:
1.3K
被引数:
1.5W
机构
引用论文
Changes in cancer detection and false-positive recall in mammography using artificial intelligence: a retrospective, multireader study使用人工智能的乳房x线照相术中癌症检测和假阳性回忆的变化: 一项回顾性多读者研究
LANCET DIGITAL HEALTH
IF24.1
AppendiXNet: Deep Learning for Diagnosis of Appendicitis from A Small Dataset of CT Exams Using Video Pretraining附录网: 使用视频预训练从ct检查的小数据集进行阑尾炎诊断的深度学习
SCIENTIFIC REPORTS
IF3.9
Subspheroids in the lithic assemblage of Barranco León (Spain): Recognizing the late Oldowan in Europe
PLOS ONE
IF0
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead停止解释高风险决策的黑盒机器学习模型,而改用可解释的模型
Opening the black box of machine learning in radiology: can the proximity of annotated cases be a way?在放射学中打开机器学习的黑匣子: 注释病例的接近可以成为一种方法吗?

