arrow
Return

A Joint-Training Two-Stage Method For Remote Sensing Image Captioning

delete2022-01-01
delete20
delete
OA
AI
X
Xiutiao Ye
王爽 cover
王爽 (Shuang Wang) *
Y
Yu Gu
J
Jihui Wang
R
Ruixuan Wang
B
Biao Hou
F
Fausto Giunchiglia
L
Licheng Jiao
DOI:10.1109/TGRS.2022.3224244delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Compared with remote sensing image (RSI) captioning methods based on the traditional encoder-decoder model, two-stage RSI captioning methods include an auxiliary remote sensing task to provide prior information, which enables them to generate more accurate descriptions. In previous two-stage RSI captioning methods, however, the image captioning and the auxiliary remote sensing tasks are handled separately, which is time-consuming and ignores mutual interference between tasks. To solve this problem, we propose a novel joint-training two-stage (JTTS) RSI captioning method. We use multilabel classification to provide prior information, and we design a differentiable sampling operator to replace the traditional nondifferentiable sampling operation to index the multilabel classification result. In contrast to previous two-stage RSI captioning methods, our method can implement joint training, and the joint loss allows the error of the generated description to flow into the optimization of the multilabel classification via backpropagation. Specifically, we approximate the Heaviside step function with the steep logistic function to implement a differentiable sampling operator for the multilabel classification. We propose a dynamic contrast loss function for multilabel classification tasks to ensure that a certain margin is maintained between the probabilities of the positive label and the negative label during sampling. We design an attribute-guided decoder to filter the multilabel prior information obtained by the sampling operator to generate more accurate image captions. The results of extensive experiments show that the JTTS method achieves state-of-the-art performance on the RSI captioning dataset (RSICD), the University of California, Merced (UCM)-captions, and the Sydney-captions datasets.
Keywords:
Image captioning
image understanding
joint training
multilabel attributes
remote sensing image (RSI)

Journal

IEEE Transactions on Geoscience and Remote Sensing cover
IEEE Transactions on Geoscience and Remote Sensing
IF:
8.6
Papers:
2.1W
Citations:
10.7W

Organization

U
University of Trento
Scholars:
8.8K
Papers: 9.0K
Citations: 1.2W
X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K