1
Return

Explainable vision transformer framework for breast lesion classification with Grad-CAM support

delete2026-07-30
delete0
delete
OA
AI
A
Ajay Kumar
M
Mudassir Khan
G
Gunjan Mittal
I
Izhar Husain
S
Sambhavi Shukla
S
Syed Arshad Ali
A
Anu Sayal
J
Janhvi Jha
R
Riaz Ahmad Ziar *
DOI:10.1186/s40644-026-01097-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Given that breast cancer remains the most lethal malignancy among women globally, there is an urgent need for reliable and comprehensible diagnostic technology. Deep learning models have demonstrated encouraging outcomes in breast imaging; nevertheless, existing methodologies struggle to encompass global context and deliver clinically relevant justifications for their conclusions. The ExViT-Breast framework was trained and internally validated on CBIS-DDSM (n = 1,566 subjects; masses and calcifications), externally validated on INbreast, and evaluated across modalities on the BUSI ultrasound dataset. A Vision Transformer backbone with a cross-view attention fusion module integrated paired craniocaudal (CC) and mediolateral oblique (MLO) mammographic views. Four explainability methods were compared: Grad-CAM, Grad-CAM++, LayerCAM, and Transformer-LRP. Classification performance was assessed by AUC and sensitivity at one false positive per examination. Explanation quality was evaluated using faithfulness (deletion/insertion curves), clinical plausibility (pointing game against expert-defined lesion regions), and ensemble-based uncertainty estimation. Model calibration was quantified using Expected Calibration Error (ECE). ExViT-Breast achieved superior diagnostic performance with an AUC of 0.94 ± 0.02 on CBIS-DDSM, significantly outperforming conventional CNN architectures (ResNet-50: AUC = 0.89, p < 0.001) and single-view Vision Transformers (AUC = 0.91, p = 0.003). External validation on INbreast demonstrated robust generalization (AUC = 0.91 ± 0.03), while cross-modality evaluation on BUSI yielded AUC = 0.88 ± 0.04. Transformer-LRP consistently provided the highest explanation quality across all datasets, achieving faithfulness scores of 0.71 ± 0.05 and plausibility rates of 79.5% on CBIS-DDSM, significantly superior to Grad-CAM variants (p < 0.001). The ExViT-Breast framework demonstrates the potential of explainable Vision Transformers in breast imaging, combining high classification performance with clinically interpretable explanations. The integration of multi-view analysis and quantitative explainability assessment positions this approach as a promising tool for computer-aided diagnosis in breast cancer screening and detection. Clinical trial number: not applicable.
Keywords:
Vision transformer
Breast cancer
Medical imaging
Explainable AI
Grad-CAM
Multi-view analysis
Deep learning
Computer-aided diagnosis

Journal

Cancer Imaging cover
Cancer Imaging
IF:
3.5
Papers:
1.3K
Citations:
3.5K

Organization

F
Faculty of Computer Science
Scholars:
183
Papers: 98
Citations: 0
C
College of Applied Medical Sciences
Scholars:
470
Papers: 421
Citations: 1
S
School of Computer Science and Engineering
Scholars:
1.1K
Papers: 512
Citations: 2
S
school of computing science & engineering
Scholars:
2
Papers: 1
Citations: 0
D
department of cse(aiml)
Scholars:
5
Papers: 3
Citations: 0
S
school of computer science & engineering
Scholars:
5
Papers: 3
Citations: 0
F
Faculty of Business and Law
Scholars:
5
Papers: 4
Citations: 0
A
applied college tanumah
Scholars:
3
Papers: 5
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers