arrow
Return

Photograph-based visual localization via data augmentation and structure consistency

delete2026-01-01
delete0
PRE
AI
L
Lee, Zih-Ying
T
Tu, Chia-Hao
L
Lu, Eric Hsueh-Chan *
DOI:10.1080/17489725.2025.2612012delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the development of Global Navigation Satellite Systems (GNSS), acquiring precise spatial coordinates for localization has become integral to modern life. However, in GNSS-denied environments such as indoor or underground spaces, alternative approaches like photograph-based visual localization are essential. In recent years, Convolutional Neural Networks (CNNs) have enabled real-time visual localization. Among CNN-based methods, pose regression approaches such as PoseNet, which directly estimate location coordinates and viewing angles, are computationally efficient but typically require large labeled datasets to achieve high accuracy. This poses a major challenge in data-scarce scenarios. To address this issue, we propose a data augmentation framework that leverages both synthetic images and CycleGAN-generated photographs to enrich the training set of PoseNet. In addition, we introduce a structure-consistent loss function to encourage the model to focus on the geometric structure of the scene rather than superficial texture patterns. Taking PoseNet as the baseline, experimental results on a benchmark outdoor dataset referencing a metric coordinate system demonstrate that our approach reduces the median position error from 1.28m to 0.89m with augmented data, and further to 0.72m with the proposed loss function, achieving a 44% improvement in positioning accuracy over the baseline.
Keywords:
Visual localization
camera pose regression
image-to-image translation
data augmentation
structure consistency

Journal

J
Journal of Location Based Services
IF:
1.4
Papers:
14
Citations:
0

Organization

N
national cheng kung university
Scholars:
3.4K
Papers: 1.4K
Citations: 0