Return
Explainable fine-grained visual classification via structured semantic representation learning for construction machinery
DOI:10.1016/j.knosys.2025.114731.png)
Abstract
En 中文
Image classification is a foundational task in computer vision. In the field of construction engineering, existing approaches have predominantly focused on coarse-grained scene-level classification, which lacks the capacity to distinguish between visually similar machinery categories such as machinery from different brands and models. In this study, a vision-based deep learning method was proposed for fine-grained visual classification of construction machinery. The technique improves image diversity through hybrid data augmentation, strengthens feature extraction for subtle model differences through feature selectors, and increases recognition accuracy via enhanced contrastive loss. We create a dataset of 2553 images, covering 8 types of construction machinery, including excavators and dump trucks from multiple brands and models to validate the method. The results show that the proposed method achieves an average accuracy of 85.90 % in classifying visually similar machinery from different brands and models. Additionally, we reveal the subtle features that drive model decisions by visualizing the attention heatmaps. Meanwhile, we developed a video analysis system, which has been validated in four key aspects of fine-grained construction management: ensuring task quality, enhancing operational safety, enabling green equipment management, and supporting sustainable maintenance.

