Return
Domain knowledge-driven image captioning for bridge damage description generation
C
Y
G
J
H
DOI:10.1016/j.autcon.2025.106116.png)
Abstract
En 中文
Deep learning-based bridge visual inspection often produces limited outputs, lacking the accurate descriptions required for practical assessments. Researchers have explored multimodal approaches to generate damage descriptions, but existing models are prone to hallucination and face challenges related to feature representation sufficiency, attention mechanism flexibility, and domain-specific knowledge integration. This paper develops an image captioning framework driven by domain knowledge to address these issues. It incorporates a multi-level feature fusion module that adaptively integrates Faster R-CNN trained weights (domain knowledge) with a CNN architecture. Additionally, it introduces a correlation-aware attention mechanism to dynamically capture interdependencies between image regions and optimise the attentional focus during LSTM decoding. Experimental results show that the proposed framework achieves higher BLEU scores and improves image-text alignment as verified through attention heatmaps. While the framework enhances inspection efficiency and quality, further dataset expansion and broader domain validation are required to assess its generalisation ability.
Keywords:
Deep learning
Image captioning
Bridge damage
Multimodal learning
Journal
IF:
11.5
Papers:
6.1K
Citations:
4.2W
Organization
No organization information available
