1
Return

SKYDET: An End-to-End Multiscale Attentive Detection Network From Foundation Models for Small Objects in Remote Sensing Images

delete2026-07-24
delete0
PRE
AI
Y
Yao Zhang
W
Wei Guo
B
Boxiang Xie
L
Lingfeng Lin
J
Jie Zhang
H
Hui Yang
Y
Yuke Meng
刘异 cover
刘异 (Yi Liu)
W
Wei Zhang
DOI:10.1109/tgrs.2026.3716766delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Remote sensing object detection remains a highly challenging task due to drastic variations in target scale, dense distributions of small objects, and complex background interference. Although existing detectors have achieved notable progress, they heavily rely on large-scale supervised pretraining and often suffer from substantial domain gaps and expensive annotation costs. In recent years, vision foundation models (VFMs) have demonstrated strong general-purpose representation capability, yet their potential in remote sensing imagery has not been fully explored. To bridge this gap, we propose SKYDET, an end-to-end robust object detection framework that explicitly migrates the billion-parameter DINOv3 model to the aerial domain. To effectively bridge the domain gap and prevent representation manifold degradation, we freeze the pretrained foundation model and propose a semantic guiding adapter (SGA) that acts as a precise semantic filter to suppress irrelevant background clutter. In addition, to address feature misalignment and semantic ambiguity during cross-scale fusion, we introduce a cross-fused encoder (CFE), whose core component is the reciprocal guidance module (RGM). The RGM establishes a reciprocal enhancement mechanism that enables spatial structure and channel semantics to guide each other, thereby effectively suppressing background noise and strengthening responses to small objects. Extensive experiments on three challenging benchmark datasets—DOTA-v1.0, AI-TOD, and NWPU VHR-10—demonstrate highly competitive performance. Specifically, SKYDET-C achieves state-of-the-art (SOTA) detection precision (with <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$AP_{50}$ </tex-math></inline-formula> scores of 72.6%, 56.1%, and 95.6%, respectively), while SKYDET-T exhibits superior advantages in high-precision localization tasks. This validates the effectiveness of transferring VFM knowledge to remote sensing tasks. Our code is available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/zhangyyy-ai-rs/SKYDET</uri>
Keywords:
Detection Transformer (DETR)
remote sensing
small-target detection
vision foundation models (VFMs)

Journal

IEEE Transactions on Geoscience and Remote Sensing cover
IEEE Transactions on Geoscience and Remote Sensing
IF:
8.6
Papers:
2.1W
Citations:
10.7W

Organization

N
ningde normal university
Scholars:
241
Papers: 89
Citations: 0
K
kth royal institute of technology
Scholars:
661
Papers: 368
Citations: 0
N
northeast forestry university
Scholars:
2.7K
Papers: 836
Citations: 0
W
wuhan university
Scholars:
7.8W
Papers: 5.7W
Citations: 70
Cited Papers

Cited Papers

Citing Papers

Citing Papers