arrow
Return

Learning consensus-aware semantic knowledge for remote sensing image captioning

delete2024-01-01
delete12
PRE
AI
Y
Yunpeng Li
X
Xiangrong Zhang *
X
Xina Cheng
X
Xu Tang
L
Licheng Jiao
DOI:10.1016/j.patcog.2023.109893delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Tremendous progresses have been made in remote sensing image captioning (RSIC) task in recent years, yet there still some unresolved problems: (1) facing the gap between the visual features and semantic concepts, (2) reasoning the higher-level relationships between semantic concepts. In this work, we focus on injecting high-level visual-semantic interaction into RSIC model. Firstly, the semantic concept extractor (SCE), end-to end trainable, precisely captures the semantic concepts contained in the RSIs. In particular, the visual-semantic co-attention (VSCA) is designed to grain coarse concept-related regions and region-related concepts for multi modal interaction. Furthermore, we incorporate the two types of attentive vectors with semantic-level relational features into a consensus exploitation (CE) block for learning cross-modal consensus-aware knowledge. The experiments on three benchmark data sets show the superiority of our approach compared with the reference methods.
Keywords:
Cross-modal understanding
Visual-semantic interaction
Remote sensing image captioning
Graph convolutional network

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

X
Xidian University
Scholars:
2.4W
Papers: 1.9W
Citations: 9.7K