返回
FusDreamer: Label-Efficient Remote Sensing World Model for Multimodal Data Classification
DOI:10.1109/TGRS.2025.3554862.png)
摘要
En 中文
World models significantly enhance hierarchical understanding, improving data integration and learning efficiency. To explore the potential of the world model in the remote sensing (RS) field, this article proposes a label-efficient RS world model for multimodal data fusion (FusDreamer). The FusDreamer uses the world model as a unified representation container to abstract common and high-level knowledge, promoting interactions across different types of data, that is, hyperspectral (HSI), light detection and ranging (LiDAR), and text data. Initially, a new latent-spatial multimodal generation (LaMG) paradigm is utilized for its exceptional information integration and detail retention capabilities. Subsequently, an open-world knowledge-guided consistency projection (OK-CP) module incorporates prompt representations for visually described objects and aligns language-visual features through contrastive learning. In this way, the domain gap can be bridged by fine-tuning the pre-trained world models with limited samples. Finally, an end-to-end multitask combinatorial optimization (MuCO) strategy can capture slight feature bias and constrain the diffusion process in a collaboratively learnable direction. Experiments conducted on four typical datasets indicate the effectiveness and advantages of the proposed FusDreamer.
Keyword:
Data models
Visualization
Data integration
Remote sensing
Laser radar
Training
Predictive models
Linguistics
Transformers
Buildings
Contrastive learning
diffusion process
multimodal data fusion
world model
期刊
IF:
8.6
论文数:
2.1W
被引数:
10.7W
机构
引用论文
暂无论文信息

