arrow
返回

Multimodal self-supervised learning for remote sensing data land cover classification

delete2025-01-01
delete0
PRE
AI
Z
Zhixiang Xue *
G
Guopeng Yang
X
Xuchu Yu
A
Anzhu Yu
Y
Yinggang Guo
B
Bing Liu
J
Jianan Zhou
DOI:10.1016/j.patcog.2024.110959delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Deep learning has revolutionized the remote sensing image processing techniques over the past few years. Nevertheless, annotating high-quality samples is difficult and time-consuming, which limits the performance of deep neural networks because of insufficient supervision information. Aiming to solve this contradiction, we investigate the multimodal self-supervised learning (MultiSSL) paradigm for pre-training and classification of remote sensing image. Specifically, the proposed self-supervised feature learning model consists of asymmetric encoder-decoder structure, in which deep unified encoder learns high-level key information characterizing multimodal remote sensing data and task-specific lightweight decoders are developed to reconstruct original data. To further enhance feature extraction capability, the cross-attention layers are utilized to exchange information contained in heterogeneous characteristics, thus learning more complementary information from multimodal remote sensing data. In fine-tuning stage, the pre-trained encoder and cross-attention layer serve as feature extractor, and leaned characteristics are combined with corresponding spectral information for land cover classification through a lightweight classifier. The self-supervised pre-training model can learn high-level key features from unlabeled samples, thereby utilizing the feature extraction capability of deep neural networks while reducing their dependence on annotated samples. Compared with existing classification paradigms, the proposed multimodal self-supervised pre-training and fine-tuning scheme achieves superior performance for remote sensing image land cover classification.
Keyword:
Remote sensing image
Unsupervised pre-training
Multimodal self-supervised
Feature learning

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

L
Lanzhou Jiaotong University
学者数:
6.3K
论文数: 3.6K
被引数: 4.2K
P
pla information engineering university
学者数:
2.8K
论文数: 1.6K
被引数: 2
N
northwest institute of nuclear technology - china
学者数:
1.3K
论文数: 659
被引数: 0
N
north china university of water resources & electric power
学者数:
4.0K
论文数: 2.8K
被引数: 1
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Structural Neural Substrates of Reading the Mind in the Eyes在眼睛中阅读心灵的结构神经基质
err2016-04-11
err0
errOAAI
errWataru Sato; Takanori Kochiyama; Shota Uono; Reiko Sawada; Yasutaka Kubota; Sayaka Yoshimura; Motomi Toichi
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容