返回
GLCM: Global-Local Captioning Model for Remote Sensing Image Captioning
DOI:10.1109/TCYB.2022.3222606.png)
摘要
En 中文
Remote sensing image captioning (RSIC), which describes a remote sensing image with a semantically related sentence, has been a cross-modal challenge between computer vision and natural language processing. For visual features extracted from remote sensing images, global features provide the complete and comprehensive visual relevance of all the words of a sentence simultaneously, while local features can emphasize the discrimination of these words individually. Therefore, not only global features are important for caption generation but also local features are meaningful for making the words more discriminative. In order to make full use of the advantages of both global and local features, in this article, we propose an attention-based global-local captioning model (GLCM) to obtain global-local visual feature representation for RSIC. Based on the proposed GLCM, the correlation of all the generated words and the relation of each separate word and the most related local visual features can be visualized in a similarity-based manner, which provides more interpretability for RSIC. In the extensive experiments, our method achieves comparable results in UCM-captions and superior results in Sydney-captions and RSICD which is the largest RSIC dataset.
Keyword:
Feature extraction
Visualization
Remote sensing
Decoding
Sensors
Natural language processing
Convolutional neural networks
Deep learning
global-local captioning model (GLCM)
image captioning
remote sensing
期刊
IF:
10.5
论文数:
1.1W
被引数:
5.0W
机构
引用论文
Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object Detection用于超高分辨率遥感目标检测的分层鲁棒卷积神经网络

