arrow
Return

A Multimodal Scale Normalization Framework for Vision-Radar Small UAV Positioning

delete2025-08-01
delete0
PRE
AI
Y
Yiyao Wan
J
Jiahuan Ji
W
Wenqing Xie
G
Guangyu Wu
F
Fuhui Zhou
吴启晖 (Qihui Wu)
DOI:10.1109/TMC.2025.3549620delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Uncrewed aerial vehicles (UAVs) positioning is of crucial importance in diverse applications. However, it is extremely challenging to realize the precise UAVs positioning over long distances due to the small size and dramatic scale variations associated with the high mobility in the wide area. To tackle this issue, a multimodal scale normalization framework is proposed for the scale-robust precise pixel-level UAV positioning. The framework exploits our proposed distance-aware image slicing and distance-aware scale normalization module. Moreover, a modal fusion-based scale normalization network is proposed that can accept arbitrary low-resolution UAV patches and produce the consistent high-resolution images at a uniform UAV instance scale with a single learnable model. The proposed framework is generic and can be directly used in the existing pixel-level positioning pipelines to improve the positioning performance and scale robustness. To verify the proposed framework in the real application, a practical vision-radar UAV positioning system is developed. Experimental results on the real-world dataset demonstrate the generality and effectiveness of our framework. Moreover, the ablation experiments also confirm the contribution of each module in the framework.
Keywords:
UAV positioning
multimodal
scale-robust
modal fusion
scale normalization

Journal

IEEE Transactions on Mobile Computing cover
IEEE Transactions on Mobile Computing
IF:
9.2
Papers:
5.6K
Citations:
1.8W

Organization

P
peking university
Scholars:
11.8W
Papers: 8.7W
Citations: 146
N
Nanjing University of Aeronautics and Astronautics
Scholars:
7.4K
Papers: 3.1K
Citations: 2.4W