Return
Multiple Conditions-Guided Diffusion Model for Remote Sensing Image Generation
DOI:10.1109/JSTARS.2026.3668715.png)
Abstract
En 中文
To address the issues of large scenes and detail attributions for generating remote sensing images (RSIs), this study proposes a multiple conditions-guided diffusion (MCGD) model. First, a prior object image control (POIC) module based on multipoint optimization, which models common ground objects, is proposed as a condition of the diffusion model. Then, a caption condition encoding and control (CCEC) module, which mines the caption semantics of the generated image, is designed to construct the semantic space and realize cross-modal transformation of features from text to image. Finally, this study proposes a visual context perception module based on attention, which deeply integrates the conditional features of POIC and CCEC to enhance the fine-grained RSI generation. Experiments show that MCGD can obtain CLIP-scores of 25.23 and 31.09, FID values of 44.26 and 10.42, and IS values of 4.6 and 7.8 on RSICD and NWPU-Captions datasets, respectively, which proves its effectiveness.
Keywords:
Semantics
Remote sensing
Feature extraction
Airports
Optimization
Diffusion models
Image synthesis
Generative adversarial networks
Noise
Microelectronics
Guided diffusion
image generation
multiple conditions
remote sensing
Journal
IF:
5.3
Papers:
1.3K
Citations:
3.0W

