arrow
Return

Semantics-Consistent Representation Learning for Remote Sensing ImageVoice Retrieval

delete2022-01-01
delete22
delete
OA
AI
H
Hailong Ning
B
Bin Zhao
Y
Yuan Yuan *
DOI:10.1109/TGRS.2021.3060705delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS imagex2013;voice retrieval provides a new insight. This article aims to study the task of RS imagex2013;voice retrieval so as to search effective information from massive amounts of RS data. Existing methods for RS imagex2013;voice retrieval rely primarily on the pairwise relationship to narrow the heterogeneous semantic gap between images and voices. However, apart from the pairwise relationship included in the data sets, the intramodality and nonpaired intermodality relationships should also be considered simultaneously since the semantic consistency among nonpaired representations plays an important role in the RS imagex2013;voice retrieval task. Inspired by this, a semantics-consistent representation learning (SCRL) method is proposed for RS imagex2013;voice retrieval. The main novelty is that the proposed method takes the pairwise, intramodality, and nonpaired intermodality relationships into account simultaneously, thereby improving the semantic consistency of the learned representations for the RS imagex2013;voice retrieval. The proposed SCRL method consists of two main steps: 1) semantics encoding and 2) SCRL. First, an image encoding network is adopted to extract high-level image features with a transfer learning strategy, and a voice encoding network with dilated convolution is devised to obtain high-level voice features. Second, a consistent representation space is conducted by modeling the three kinds of relationships to narrow the heterogeneous semantic gap and learn semantics-consistent representations across two modalities. Extensive experimental results on three challenging RS imagex2013;voice data sets, including Sydney, UCM, and RSICD imagex2013;voice data sets, show the effectiveness of the proposed method.
Keywords:
Semantics
Feature extraction
Task analysis
Image retrieval
Image coding
Space exploration
Remote sensing
Heterogeneous semantic gap
remote sensing (RS) image-voice retrieval
semantics-consistent representation
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Geoscience and Remote Sensing cover
IEEE Transactions on Geoscience and Remote Sensing
IF:
8.6
Papers:
2.1W
Citations:
10.7W

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
C
chinese academy of sciences
Scholars:
55.9W
Papers: 44.7W
Citations: 704