Return
Photograph-based visual localization via data augmentation and structure consistency
DOI:10.1080/17489725.2025.2612012.png)
Abstract
En 中文
With the development of Global Navigation Satellite Systems (GNSS), acquiring precise spatial coordinates for localization has become integral to modern life. However, in GNSS-denied environments such as indoor or underground spaces, alternative approaches like photograph-based visual localization are essential. In recent years, Convolutional Neural Networks (CNNs) have enabled real-time visual localization. Among CNN-based methods, pose regression approaches such as PoseNet, which directly estimate location coordinates and viewing angles, are computationally efficient but typically require large labeled datasets to achieve high accuracy. This poses a major challenge in data-scarce scenarios. To address this issue, we propose a data augmentation framework that leverages both synthetic images and CycleGAN-generated photographs to enrich the training set of PoseNet. In addition, we introduce a structure-consistent loss function to encourage the model to focus on the geometric structure of the scene rather than superficial texture patterns. Taking PoseNet as the baseline, experimental results on a benchmark outdoor dataset referencing a metric coordinate system demonstrate that our approach reduces the median position error from 1.28m to 0.89m with augmented data, and further to 0.72m with the proposed loss function, achieving a 44% improvement in positioning accuracy over the baseline.
Keywords:
Visual localization
camera pose regression
image-to-image translation
data augmentation
structure consistency
Journal
J
IF:
1.4
Papers:
14
Citations:
0

