Return
CSSA: A Cross-Modal Spatial–Semantic Alignment Framework for Remote Sensing Image Captioning
DOI:10.3390/rs18030522.png)
Abstract
En
Keywords:
remote sensing image captioning
multi-branch cross-modal contrastive learning
dynamic geometry transformer
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.1
Papers:
6.9K
Citations:
15.1W

