1
Return

Domain knowledge-driven image captioning for bridge damage description generation

delete2025-06-01
delete0
PRE
AI
C
Chengzhang Chai
Y
Yan Gao
G
Guanyu Xiong
J
Jiucai Liu
H
Haijiang Li *
DOI:10.1016/j.autcon.2025.106116delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep learning-based bridge visual inspection often produces limited outputs, lacking the accurate descriptions required for practical assessments. Researchers have explored multimodal approaches to generate damage descriptions, but existing models are prone to hallucination and face challenges related to feature representation sufficiency, attention mechanism flexibility, and domain-specific knowledge integration. This paper develops an image captioning framework driven by domain knowledge to address these issues. It incorporates a multi-level feature fusion module that adaptively integrates Faster R-CNN trained weights (domain knowledge) with a CNN architecture. Additionally, it introduces a correlation-aware attention mechanism to dynamically capture interdependencies between image regions and optimise the attentional focus during LSTM decoding. Experimental results show that the proposed framework achieves higher BLEU scores and improves image-text alignment as verified through attention heatmaps. While the framework enhances inspection efficiency and quality, further dataset expansion and broader domain validation are required to assess its generalisation ability.
Keywords:
Deep learning
Image captioning
Bridge damage
Multimodal learning

Journal

Automation in Construction cover
Automation in Construction
IF:
11.5
Papers:
6.1K
Citations:
4.2W

Organization

No organization information available
Cited Papers

Cited Papers

Citing Papers

Citing Papers