返回
Incorporating object counts into remote sensing image captioning
DOI:10.1080/17538947.2024.2392847.png)
摘要
En 中文
Existing methods for remote sensing image captioning tend to describe a remote sensing image using generic language that lacks specific information about object counts. To address this limitation, we propose a novel framework for generating a caption that includes object count information for the remote sensing image. Our proposed framework comprises three modules: object counting, preliminary captioning, and numeral editing. The object counting module identifies objects in a remote sensing image and determines object counts. The preliminary captioning module generates a caption that may lack object count information. The numeral editing module incorporates the object counts into the caption, resulting in a more precise caption. Our proposed framework outperforms existing methods, as demonstrated through evaluations on three remote sensing image datasets. Our proposed framework is a significant step toward more precise and informative remote sensing image captioning.
Keyword:
Remote sensing
earth observation
artificial intelligence
image processing
期刊
IF:
4.9
论文数:
2.0K
被引数:
4.7K
机构
引用论文
Hierarchical and Robust Convolutional Neural Network for Very High-Resolution Remote Sensing Object Detection用于超高分辨率遥感目标检测的分层鲁棒卷积神经网络
Meta captioning: A meta learning based remote sensing image captioning framework元字幕: 一种基于元学习的遥感图像字幕框架
Reinforcement learning of non-additive joint steganographic embedding costs with attention mechanism

