Return
Adaptive Scale-Aware Semantic Memory Network for Remote Sensing Image Captioning
DOI:10.1109/TGRS.2025.3636596.png)
Abstract
En 中文
Remote sensing image captioning between visual images and natural language remains a long-standing challenge in the remote sensing community. Due to the wide coverage and large amount of information in remote sensing images, existing methods struggle to effectively utilize the relevant semantic information about objects and their attributes at different scales across samples to generate descriptions. To address these issues, the article proposes a novel adaptive scale-aware semantic memory network (ASSMN) for remote sensing image captioning. First, to fully extract the semantic information in remote sensing images, multilevel feature enhancement is constructed to improve the feature representation extracted from the contrastive language-image pretraining (CLIP) pretraining model. Subsequently, a scale-aware attention aggregator (SAA) is introduced to further integrate the enhanced multiscale image features into the high-level semantics of remote sensing images. Then, to fully exploit the semantic information of the joint observed samples, a semantic memory reinforcement (SMR) is designed to strengthen the semantic representation of the current scene through the relevant semantics obtained from other training samples. Finally, a captioning decoder is performed to generate a comprehensive scene caption with accurate objects and attributes. In the experiments, the performance of the proposed ASSMN on three remote sensing image captioning datasets is evaluated and compared with well-designed baselines and state-of-the-art methods, and the superiority of the proposed method is demonstrated by ablating the role of each proposed component. The code will be available at https://github.com/zcsisiyao/ASSMN
Keywords:
Contrastive language image pretrain
image captioning
prototype memory
remote sensing
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

