Return
FusDreamer: Label-Efficient Remote Sensing World Model for Multimodal Data Classification
DOI:10.1109/TGRS.2025.3554862.png)
Abstract
En 中文
World models significantly enhance hierarchical understanding, improving data integration and learning efficiency. To explore the potential of the world model in the remote sensing (RS) field, this article proposes a label-efficient RS world model for multimodal data fusion (FusDreamer). The FusDreamer uses the world model as a unified representation container to abstract common and high-level knowledge, promoting interactions across different types of data, that is, hyperspectral (HSI), light detection and ranging (LiDAR), and text data. Initially, a new latent-spatial multimodal generation (LaMG) paradigm is utilized for its exceptional information integration and detail retention capabilities. Subsequently, an open-world knowledge-guided consistency projection (OK-CP) module incorporates prompt representations for visually described objects and aligns language-visual features through contrastive learning. In this way, the domain gap can be bridged by fine-tuning the pre-trained world models with limited samples. Finally, an end-to-end multitask combinatorial optimization (MuCO) strategy can capture slight feature bias and constrain the diffusion process in a collaboratively learnable direction. Experiments conducted on four typical datasets indicate the effectiveness and advantages of the proposed FusDreamer.
Keywords:
Data models
Visualization
Data integration
Remote sensing
Laser radar
Training
Predictive models
Linguistics
Transformers
Buildings
Contrastive learning
diffusion process
multimodal data fusion
world model
Journal
IF:
8.6
Papers:
2.1W
Citations:
10.7W

