返回
Multimodal self-supervised learning for remote sensing data land cover classification
DOI:10.1016/j.patcog.2024.110959.png)
摘要
En 中文
Deep learning has revolutionized the remote sensing image processing techniques over the past few years. Nevertheless, annotating high-quality samples is difficult and time-consuming, which limits the performance of deep neural networks because of insufficient supervision information. Aiming to solve this contradiction, we investigate the multimodal self-supervised learning (MultiSSL) paradigm for pre-training and classification of remote sensing image. Specifically, the proposed self-supervised feature learning model consists of asymmetric encoder-decoder structure, in which deep unified encoder learns high-level key information characterizing multimodal remote sensing data and task-specific lightweight decoders are developed to reconstruct original data. To further enhance feature extraction capability, the cross-attention layers are utilized to exchange information contained in heterogeneous characteristics, thus learning more complementary information from multimodal remote sensing data. In fine-tuning stage, the pre-trained encoder and cross-attention layer serve as feature extractor, and leaned characteristics are combined with corresponding spectral information for land cover classification through a lightweight classifier. The self-supervised pre-training model can learn high-level key features from unlabeled samples, thereby utilizing the feature extraction capability of deep neural networks while reducing their dependence on annotated samples. Compared with existing classification paradigms, the proposed multimodal self-supervised pre-training and fine-tuning scheme achieves superior performance for remote sensing image land cover classification.
Keyword:
Remote sensing image
Unsupervised pre-training
Multimodal self-supervised
Feature learning
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Deep support vector machine for hyperspectral image classification基于深度支持向量机的高光谱图像分类
PATTERN RECOGNITION
IF7.6
Hyperspectral and SAR Image Classification via Multiscale Interactive Fusion Network基于多尺度交互式融合网络的高光谱与SAR图像分类
Joint bilateral filtering and spectral similarity-based sparse representation: A generic framework for effective feature extraction and data classification in hyperspectral imaging联合双边滤波和基于光谱相似性的稀疏表示: 高光谱成像中有效特征提取和数据分类的通用框架
PATTERN RECOGNITION
IF7.6
Unsupervised Spatial-Spectral Feature Learning by 3D Convolutional Autoencoder for Hyperspectral Classification基于3D卷积自动编码器的无监督空间光谱特征学习,用于高光谱分类
Multimodal remote sensing benchmark datasets for land cover classification with a shared and specific feature learning model具有共享和特定特征学习模型的多模态遥感基准数据集,用于土地覆盖分类

