Return
Relation-Prior-Guided Asymmetric Focal Loss for Open-Vocabulary Remote Sensing Object Detection
J
M
B
X
DOI:10.1109/lgrs.2026.3711619.png)
Abstract
En 中文
Open-vocabulary object detection (OVD) achieves strong zero-shot transferability in natural images by aligning region features with text embeddings. However, this transferability largely breaks down in remote sensing (RS) images, showing significant drops in both novel-class recall and average precision. We attribute this degradation to a text–visual inconsistency in category relations: categories with similar overhead-view appearances are not necessarily close in the text embedding space, making text embeddings less reliable transfer anchors. To address this, we construct an RS relation prior (RSRP) that jointly models text semantic similarity, RS visual similarity, and text–visual consistency, without using any novel-class annotations. Building upon RSRP, we further propose a relation-prior-guided asymmetric focal loss (RPA-FL), which relaxes excessive suppression on visually related base classes and provides weak positive supervision (WPS) for novel classes with low text–visual consistency. Experiments on DIOR and DOTA demonstrate that the proposed method outperforms state-of-the-art RS OVD approaches under both zero-shot and generalized zero-shot detection protocols, without additional inference overhead. The code and configurations are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/perchatao/RPA-FL</uri>
Keywords:
Focal loss
Grounding DINO
open-vocabulary detection
relation prior
remote sensing (RS) images
Journal
I
IF:
4.4
Papers:
486
Citations:
0
