arrow
Return

Inspector gaze-guided multitask learning for explainable structural damage assessment

delete2025-11-01
delete0
PRE
AI
C
Chenyu Zhang
L
Liu, Charlotte
K
Ke Li
Z
Zhaozheng Yin
DOI:10.1111/mice.70131delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Accurately classifying damage levels from structural inspection images is critical for automated infrastructure assessment. Although deep neural networks achieve impressive performance, their black-box nature limits explainability, and prior studies using Grad-CAM often yield coarse or inaccurate saliency maps. To overcome these limitations, this paper introduces XIDLE-Net, a multitask model that simultaneously performs damage classification and saliency map prediction to enhance explainability in structural damage assessment. Combining a Swin Transformer encoder with a convolutional neural network decoder, XIDLE-Net is trained with dual supervision using damage labels and inspector gaze-derived attention maps, enhancing both classification accuracy and model explainability. Experimental results show that XIDLE-Net outperforms state-of-the-art methods in both classification and saliency explainability, achieving 78.1% accuracy, 94.3% area under the curve (AUC), and a 39.7% improvement in saliency prediction over ResNet-50 with Grad-CAM. To our knowledge, this is one of the first investigations to employ large-scale inspector gaze data for supervision and to quantitatively evaluate Grad-CAM in structural image classification. The results highlight the promise of human gaze data for advancing explainable vision-based structural health monitoring.
Keywords:
CONVOLUTIONAL NEURAL-NETWORKS
DEEP
CLASSIFICATION

Journal

C
Computer-Aided Civil and Infrastructure Engineering
IF:
9.1
Papers:
2.0K
Citations:
10.0K

Organization

S
state university of new york (suny) system
Scholars:
6.5W
Papers: 5.8W
Citations: 65