Return
Interactive Concept Network Enhanced Transformer for Remote Sensing Image Captioning
DOI:10.1109/TGRS.2024.3523305.png)
Abstract
En 中文
Remote sensing image captioning plays an important role in advancing remote sensing image understanding with natural language generation. However, it is difficult to generate accurate semantic descriptions of crucial objects and their relationships, due to large coverage and abundant information in remote sensing images. To address these issues, this article proposes a novel interactive concept network enhanced transformer (ICNET) for remote sensing image captioning. First, multilevel visual features are extracted within a local and global feature extraction module. To comprehensively capture key objects in the local features, a concept mapping network (CMN) is constructed to project multiscale local features onto high-level semantic concepts of the objects. This allows for the integration of the relevant feature vectors in the visual feature mapping into multiple relatively independent word features, thus bridging the gap between visual features and semantic concepts. Subsequently, a global feature enhancement (GFE) module is introduced to boost the discrimination of global relationships and filter irrelevant content. Finally, to aggregate semantic concepts and global features, a transformer equipped with a concept interaction module (CIM) is designed to facilitate feature alignment and generate captions with proper categories and relationships. The experimental results on three remote sensing image captioning datasets demonstrate the superiority of the proposed method.
Keywords:
Feature extraction
Visualization
Transformers
Semantics
Vectors
Convolutional neural networks
Aggregates
Data mining
Computer vision
Concept network
image captioning
relational description
remote sensing
transformer
transformer
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

