Return
Semantics-Consistent Representation Learning for Remote Sensing ImageVoice Retrieval
DOI:10.1109/TGRS.2021.3060705.png)
Abstract
En 中文
With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS imagex2013;voice retrieval provides a new insight. This article aims to study the task of RS imagex2013;voice retrieval so as to search effective information from massive amounts of RS data. Existing methods for RS imagex2013;voice retrieval rely primarily on the pairwise relationship to narrow the heterogeneous semantic gap between images and voices. However, apart from the pairwise relationship included in the data sets, the intramodality and nonpaired intermodality relationships should also be considered simultaneously since the semantic consistency among nonpaired representations plays an important role in the RS imagex2013;voice retrieval task. Inspired by this, a semantics-consistent representation learning (SCRL) method is proposed for RS imagex2013;voice retrieval. The main novelty is that the proposed method takes the pairwise, intramodality, and nonpaired intermodality relationships into account simultaneously, thereby improving the semantic consistency of the learned representations for the RS imagex2013;voice retrieval. The proposed SCRL method consists of two main steps: 1) semantics encoding and 2) SCRL. First, an image encoding network is adopted to extract high-level image features with a transfer learning strategy, and a voice encoding network with dilated convolution is devised to obtain high-level voice features. Second, a consistent representation space is conducted by modeling the three kinds of relationships to narrow the heterogeneous semantic gap and learn semantics-consistent representations across two modalities. Extensive experimental results on three challenging RS imagex2013;voice data sets, including Sydney, UCM, and RSICD imagex2013;voice data sets, show the effectiveness of the proposed method.
Keywords:
Semantics
Feature extraction
Task analysis
Image retrieval
Image coding
Space exploration
Remote sensing
Heterogeneous semantic gap
remote sensing (RS) image-voice retrieval
semantics-consistent representation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

