Return
TG-DANet: Text-Guided Dual-Awareness Network for Oriented Object Detection
W
H
P
Z
H
DOI:10.1142/S0218001426550062.png)
Abstract
En 中文
Oriented object detection (OOD) has rapidly advanced in recent years. However, the performance of existing methods is unsatisfactory when dealing with challenging scenarios, especially in scenes involving small-scale objects or objects with extreme aspect ratio. Inspired by recent advances in vision-language pre-training, we propose a novel Text-Guided Dual-Awareness Network (TG-DANet), which addresses these challenges from two complementary perspectives: robust feature interaction for multi-scale and long-range context modeling, and semantic-aware feature learning through textual guidance. Specifically, we design a Bi-Directional Feature Interaction Module (BDFIM) to capture horizontal and vertical contextual features via spatial interactions, which improves the representation of small and elongated objects. Additionally, a Text-Semantic Guidance Framework (TSGF) is supposed to align and fuse textual embeddings with visual features at multiple levels, which enhances model interpretability and discriminability for objects with ambiguous appearances or complex layouts. Extensive experiments on three benchmark datasets (DOTA, DIOR-R, and HRSC2016) show that TG-DANet achieves improvements of 3.05%, 3.49%, and 2.32% in mAP over baseline methods, respectively. These results demonstrate the effectiveness of our dual-perspective strategy in handling complex scenes with cluttered backgrounds and multi-scale objects, which highlights the promising potential of vision-language fusion in OOD.
Keywords:
Remote sensing
oriented object detection
bi-directional feature interaction
textual semantic guidance
dual-awareness network
Journal
IF:
1.1
Papers:
161
Citations:
2.0K
