Return
Could Describe Anything Model Understand Remote Sensing Objects?
DOI:10.1109/JSTARS.2025.3616330.png)
Abstract
En 中文
Fine-grainedimage captioning for remote sensing images (RSIs) remains an underexplored yet crucial task for multimodal scene understanding, especially in domains where localized semantic interpretation is essential. However, the heterogeneous nature of remote sensing data—spanning modalities, such as visible light, infrared, and synthetic aperture radar (SAR)—poses significant challenges for generalization, particularly under the lack of sentence-level annotations. Existing vision-language models, primarily trained on natural image–text pairs, often fail to capture spatially grounded semantics in RSIs, leading to degraded performance when transferred to nonnatural domains. In this article, we present the first systematic evaluation of the describe anything model (DAM) for localized captioning in remote sensing. To address the scarcity of aligned supervision, we construct a weakly supervised benchmark framework featuring three levels of region-of-interest prompting (full image, center point, and bounding box), and harmonize test data across three modalities with a unified object category (ship). Caption quality is assessed through both linguistic diversity metrics and multimodal alignment indicators, including contrastive language-image pretraining (CLIP) and RemoteCLIP scores. Experimental results reveal that DAM performs robustly on visible imagery with fine-grained prompts, but exhibits significant performance degradation in infrared and SAR domains, where modality-specific distortions hinder effective spatial grounding. Our benchmark exposes a critical bottleneck in current foundation models for cross-modality captioning and establishes a unified testbed for evaluating and improving multimodal language understanding in the remote sensing field.
Keywords:
Describe anything model (DAM)
localized image captioning
remote sensing image (RSI)
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
5.3
Papers:
1.3K
Citations:
3.0W

