Return
DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization
DOI:10.1109/JIOT.2026.3661248.png)
Abstract
En 中文
With the rapid growth of the low-altitude economy, uncrewed aerial vehicles (UAVs) have become key platforms for measurement and tracking in intelligent patrol systems. However, in global navigation satellite system (GNSS)-denied environments, localization schemes that rely solely on satellite signals are prone to failure. Cross-view image retrieval-based localization is a promising alternative, yet substantial geometric and appearance domain gaps exist between oblique UAV views and nadir satellite orthophotos. Moreover, conventional approaches often depend on complex network architectures, text prompts, or large amounts of annotation, which hinders generalization. To address these issues, we propose DiffusionUavLoc, a cross-view localization framework that is image-prompted, text-free, diffusion-centric, and employs a variational autoencoder (VAE) for unified representation. We first use training-free geometric rendering to synthesize pseudo-satellite images from UAV imagery as structural prompts. We then design a text-free conditional diffusion model that fuses multimodal structural cues to learn features robust to viewpoint changes. Inference, descriptors are computed at a fixed timestep $t$ and compared using cosine similarity. On University-1652 and SUES-200, the method performs competitively for cross-view localization, especially for satellite-to-drone in University-1652. Our data and code will be published at the following URL: https://github.com/liutao23/DiffusionUavLoc.git
Keywords:
ControlNet
cross-view image retrieval
diffusion models
uncrewed aerial vehicle (UAV) localization
variational autoencoder (VAE)
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

