Return
Bridge damage description using adaptive attention-based image captioning
DOI:10.1016/j.autcon.2024.105525.png)
Abstract
En 中文
Current many vision -based research uses various classification, detection, and segmentation methods to identify bridge damage. Instead of these numerical results, a highly abstract natural language description is a more suitable method to summarize and transmit bridge inspection processes and results to humans. This paper presents an end -to -end image captioning -based bridge damage comprehensive description network (BDCD-Net) for describing and locating bridge damage. BDCD-Net consists of two parts: an image feature encoder (extracting multi -level image features from bridge damage images) and a bridge damage description generation decoder (employing an adaptive attention mechanism to selectively utilize image features to generate descriptions and locate damage). The descriptions include component types, damage categories, relative spatial positions of the damaged components and bridges, and shooting angles of the image. The effectiveness of BDCD-Net was validated using images collected from real bridges with annotated descriptions. The results indicate the significant potential of fully automated bridge inspection.
Keywords:
Bridge damage
Comprehensive description
Image captioning
Adaptive attention
Multimodal learning
Journal
IF:
11.5
Papers:
6.2K
Citations:
4.2W

