Return
Text-Augmented Semantic Feature Extraction and Difference Information Learning for Remote Sensing Image Change Captioning
DOI:10.1109/TGRS.2025.3595422.png)
Abstract
En 中文
Remote sensing image change captioning (RSICC) aims to generate sentence descriptions about land cover changes in bitemporal images. The effective acquisition of semantic-level change information is critical for this task. However, due to the effects of illumination interference, appearance similarities and scale differences between different objects, it is difficult to accurately extract change information from bitemporal images. In this article, we attempt to take advantage of the high-level semantic information inherent in text and propose a text-augmented semantic feature extraction and difference information learning model for RSICC. Specifically, we first predefine some text prompts for each remote sensing image and use the contrastive language-image pretraining (CLIP) model to select the most suitable text descriptions for them. Then, we adopt a refined segment anything model (SAM) to learn fine-grained visual features from each image, which is further enhanced via a designed selective text–image fusion (STIF) module. After that, to extract the semantic differences between bitemporal images, we propose a text-guided difference capture (TGDC) module capable of extracting multiscale difference information under the guidance of text differences between different-time images. Finally, a transformer-based caption generator is applied to generate sentence descriptions from the extracted difference information. In order to test the performance of our proposed model, we conduct comprehensive experiments on two widely used RSICC datasets, including LEVIR-CC and Dubai-CC. The experimental results show that our proposed model is able to outperform several state-of-the-art models, which validates the effectiveness of it. The codes of our proposed model will be released at https://github.com/Richardkimyo/TACC
Keywords:
Change captioning
difference information learning
semantic feature extraction
text prompts
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

